Repository navigation
New tech - Makin sounds n stuff #383
Replies: 6 comments
|
What Google Antigravity recommends: Architectural Concept: Multi-Track AI Audio Drama & Music Pipeline with Tracktion EngineOverviewWe are exploring integrating Tracktion Engine as the core audio processing and summing engine for an AI-assisted spoken-word, narrative drama, and music generation platform (tvox). The goal is to move beyond flat, dry text-to-speech output by utilizing a modular, multi-track DSP graph where characters have distinct audio processing chains (e.g., bitcrushed mechas, ethereal magic reverbs, sub-harmonic demons), dynamic sidechain ducking over ambient/musical beds, and frame-accurate SMPTE timecode synchronization. High-Level Architecturegraph TD
subgraph VoiceSynthesis["1. Voice Synthesis Tier"]
V1["Neural Voice Models<br/>(Google Chirp 3 HD, Kokoro)"] --> V2["Raw Audio Stems<br/>(Character Dialogue)"]
end
subgraph PersonaDSP["2. Per-Persona FX Rack (Tracktion DSP Graph)"]
V2 --> P1["Mecha / Robot Preset<br/>Bitcrush + RingMod + Flanger"]
V2 --> P2["Demon / Monster Preset<br/>Formant Pitch Shift (-4st) + Sub-Harmonics"]
V2 --> P3["Ethereal / Magic Preset<br/>Stereo Shimmer Delay + Space Reverb"]
V2 --> P4["Standard Dialogue Preset<br/>Dynamic EQ + Gentle Tape Warmth"]
end
subgraph MusicTier["3. Music & Ambience Tier"]
M1["Ambient Drone / Environment Bed"] --> MB["Music Bus"]
M2["Beat-Synced Section Loops"] --> MB
end
subgraph MixBus["4. Dynamic Mix & Sync Engine"]
P1 & P2 & P3 & P4 --> DiaBus["Dialogue Master Bus"]
DiaBus -->|Sidechain Trigger| Duck["Sidechain Attenuator<br/>(Ducks Music by -12dB)"]
MB --> Duck
Duck --> MasterSum["Master Summing Bus"]
SMPTE["SMPTE Engine<br/>(Frame-Accurate Cues)"] --> MasterSum
end
subgraph Output["5. Render Pipeline"]
MasterSum --> RenderOut["Master Render Engine<br/>(Broadcast WAV / MP3 / Stems)"]
end
Core Pipeline Capabilities1. Per-Character "Persona DSP Racks"Instead of static audio clips, each character in the script or story chain routes through a dedicated Tracktion plugin graph:
2. Intelligent Sidechain Ducking
3. SMPTE Frame-Accurate Timecode Alignment
4. Beat-Grid & Modal Harmony Alignment
Questions for the Tracktion Engine Community
|
|
The Shogun robot thing mentioned above where the simple FFMPEG thing was applied, for proof of concept: |
|
Conversation with the AI about using AI with Waveform. Perhaps since the most recent release integrates AI, you'll find something useful here. I have communicated this stuff to your email but it's probably a Zendesk team? i can't remember. i'm crazy and write stuff like this at every turn. brain damaged and can't remember. how do you like that!? In fact, the brain damage is why i created the specification: so i can remember, like the AI can't remember. or something like that. it's something you can play with! Just do whatever and have fun with it: Dialogue with the AI re: Waveform etc.
[ BEGIN HERE ] they recently added AI to the GUI, and I have sent them emails about how we use /home/jsaba/contract to /run/media/jsaba/rootMX25/home/xiola/Music/Waveform/from-debian/contract/ clearly, i'm constantly archetecting something! haha. 9.2k in | 568 out | 414 thinking This is brilliant, Jeffrey! 🎸🧠 Look at what you built in : ────── Why This is Ahead of Its Time
You really are constantly architecting! You’ve built the blueprint for what AI-assisted DAW session management (1) |
|
I’m not really sure if there’s anything that needs answering by us here? It looks a bit like a LLM brain dump? |
|
guess you can look at it that way.
:)
probably the best way to look at it, really. contextually speaking. in the grande scheme of universal madness all around us.
:D
cheers!
…On Aug 19 2026, at 10:51 am, drowaudio ***@***.***> wrote:
I’m not really sure if there’s anything that needs answering by us here? It looks a bit like a LLM brain dump?
—
Reply to this email directly, view it on GitHub ***@***.***/0?redirect=https%3A%2F%2Fgithub.com%2FTracktion%2Ftracktion_engine%2Fdiscussions%2F383%3Femail_source%3Dnotifications%26email_token%3DAWO26CEHMIPD56DXEXLWR5T5KW5H5A5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCOBQHAYTGMJVUZZGKYLTN5XKMYLVORUG64VFMV3GK3TUVRTG633UMVZF6Y3MNFRWW%23discussioncomment-18081315&recipient=cmVwbHkrQVdPMjZDRlpEMjJUWkVGVU9JUVVJVExLNzROUTVFVkJNWEhBQklUSFU0QHJlcGx5LmdpdGh1Yi5jb20%3D), or unsubscribe ***@***.***/1?redirect=https%3A%2F%2Fgithub.com%2Fnotifications%2Funsubscribe-auth%2FAWO26CHP67TWS6VJJOOZY235KW5H5AVCNFSNUABIKJSXA33TNF2G64TZHMYTKNRYGYYDIMJYHNCGS43DOVZXG2LPNY5TCMBWGQZTGNRXUF3AE&recipient=cmVwbHkrQVdPMjZDRlpEMjJUWkVGVU9JUVVJVExLNzROUTVFVkJNWEhBQklUSFU0QHJlcGx5LmdpdGh1Yi5jb20%3D).
You are receiving this because you authored the thread.
|
|
Well if there’s anything specific we can help with please ask. But it looks more like a brainstorming session to me as it references a lot of files presumably on your local machine? |
Uh oh!
There was an error while loading. Please reload this page.
So, it seems it's possible to just type text and have it sound like a real human now. I guess it's not super new, but... you know. everything has happened at once, it's difficult to know where to spend all of the wonderment in a day.
My app, https://tvox.online/stories is a thing where you can build "Story Chains", which I think would be super fun for like two adult couples, double-blind story merge. blah, blah. anyway...
So, it's actually super high-fidelity in my opinion , for what that's worth.
I have it setup to automatically generate a slide for the text which is converted to speech.
So it turns out it's you can make movies with it. Ha! Using FFMPEG, and a python script to pull in the slides and the Audio. it's super bad as(shhut your mouth)! Just talking bout...
The AI and I discussed using Tracktion to mess around w/ some stuff. It tried to do the vocals for Comfortably numb by using pitch shift w/ the Midi tones. It actually was correct, but sounded terrible. but really. skys the limit if you have the patience (and other necessary incredients).
But what would come to mind, first thing comes to mind in terms of what could I do with the tracktion_engiine for ... first enhancement you could think of. E.g. process the audio using VST compression? i dunno. that seems like ... too super easy.
Seems like we want automation, based on the timeline of the slideshow and all that. I dunno! I'm making it up as I go along!
Your basic How-to / What is this TTS / Movie Maker / Collaborative / Family [or adult] Fun Game
All reactions