Video summary
How I Built an End-to-End Local Voice Agent, and Made It Fast
Main summary
Key takeaways
End-to-end local voice agent (local “AI stack”)
- Built a fully local voice agent where the entire pipeline runs on a single 12GB GPU.
- Uses three core model components:
- LLM: ~35B Mixture-of-Experts (Qwen3.5 / Qwen 3.6 35B MoE) for “real work,” not just chat.
- Speech-to-text: Whisper for listening.
- Text-to-speech / voice: Breeze TTS2 (3.5B) as the talking voice.
- Adds turn detection for hands-free operation using Cilero VAD (detects when the user starts/stops talking).
Motivation: prior experiments had multi-second latency (5–6s) and sometimes required two machines; this build targets “feels like talking to a real person.”
Key technical goal: reduce latency (and the “silence” problem)
- The hard part isn’t wiring models—it’s making the system fast enough and responsive enough to avoid awkward waiting.
- Measured baseline: about 10–11 seconds from user stop talking to first response audio.
- It gets worse for longer answers because:
- The system waits for the model to finish writing the whole response before the voice begins speaking.
Performance & responsiveness optimizations (multi-part strategy)
1) Cut long initial silence (streaming response by sentence/phrase)
- Fix: don’t wait for the entire answer.
- As the model generates text, the text is chunked into short phrases (e.g., ~3 words) and sent to the TTS engine progressively.
- This reduces “long silence while writing.”
2) Reduce pauses between sentences (parallelize TTS generation + playback)
- Fix: while a sentence is being spoken, the system starts generating the next sentence’s audio in advance.
- Uses a queue so the speaker always has audio ready.
- Remaining issue: if the next sentence is very long, the queue can sometimes empty briefly.
3) Improve initial “time-to-first-sound”
- Baseline voice generation speed was near the edge: voice audio generation ~1.1× real time, which can cause occasional gaps.
- Further fix targets other pipeline waits.
Pipeline-level latency reductions (transcription + thinking behavior)
4) Transcribe while the user is still talking (not after)
- Previously: waited for speech end → then transcribed the entire recording.
- Fix: start transcription in parallel every couple seconds while the user speaks.
- After stop, it still keeps about ~1 second of silence before responding to avoid triggering mid-pause interruptions.
5) “Thinking model” adjustment (faster first response)
- The Qwen model can “think” before speaking.
- Fix: disable thinking only for the first response after the user speaks.
- Later parts (tool calls, actual work) can still allow thinking for quality.
Fit everything into 12GB (practical VRAM accounting + quantization)
- Relies on careful memory budgeting:
- Context length tradeoffs: reduced from larger context sizes (e.g., 128k → 64k → 32k) to fit in VRAM.
- Voice model dominates GPU usage, requiring context reduction.
- Major speed/memory win:
- Breeze TTS2 switched from BF16 to Q8 quantized (via
audio.cpp), resulting in:- 2–3× faster generation
- ~4GB VRAM freed
audio.cppalso supports streaming audio, reducing speaker waiting for full sentence completion.
- Breeze TTS2 switched from BF16 to Q8 quantized (via
Agent tooling upgrades: “real work” + visible UI
6) Agent harness: Pythagoras
- All changes are integrated into the speaker’s agent framework called Pythagoras.
- Includes UI “work panels” for debugging/visibility:
- Browser panel
- Terminal panel
- Canvas
- Displays when the AI uses tools (with transitions/animations)
7) Faster tool use: speed up prefill + long tool outputs
- Tool read cost problem: web pages can be tens of thousands of tokens.
- Observed: one YouTube page took ~1.5 minutes to read.
- Cause: micro batch size 128 leading to slow prefill.
Fixes:
- Free VRAM via voice quantization (Q8).
- Increase micro batch size to ~1024.
- Pull expert layers back onto GPU.
- Restore context range upward (eventually up to ~100,000 tokens), while keeping spare space for spikes.
- Added efficiency: conversation caching so it doesn’t reread everything every turn; cache persisted to disk and reloaded.
8) Talking while working (avoid silence during tool calls/thinking)
- New behavior in voice mode:
- User messages get a tag (e.g., indicating they came from the voice interface).
- System prompt makes voice replies brief.
- Before tool use, the AI must say out loud what it will do.
- Additional fix for “silent thinking/output reading”:
- When the system starts thinking or waiting on work, it sends pre-chosen ready-made phrases (randomized) like “Let me think about that…”
- Also handles context overflow:
- Announces when context is full and it will compact the conversation.
End results (measured improvements)
- Turn latency improved from ~8 seconds into the ~2–3 second range across turns.
- Faster agent behavior in practice:
- Example: open YouTube → search channel → describe first image detected on the page.
- Example: switch to local llama subreddit → generate a report summarizing what’s trending (tool+web workflow).
Build philosophy / extensibility
- The interface and orchestration are a major part of “feels like the future,” not only raw intelligence.
- Modular approach:
- The model is the main swappable component, and improvements to local models can be dropped in.
- Configuration is portable:
- Numbers depend on your card, but the strategy (context/compute buffering/voice quantization) stays the same.
- Code is available as open source: integrated in Pythagoras on GitHub.
- Encourages using it and submitting PRs for gaps.
Main speakers / sources
- Speaker: the video creator (self-referential “I” throughout; presenting measurements and building steps).
- External sources mentioned:
- OpenAI (reference to the Astra demo as inspiration)
- Qwen (Qwen3.5 / Qwen MoE models referenced)
- Whisper (speech-to-text)
- Breeze TTS2 (voice model; referenced as top in an “Artificial analysis leaderboard”)
- Cilero VAD (turn detection)
- CodeRabbit (sponsor; PR reviewer tool used to catch race conditions)
- GitHub (Pythagoras open-source repository)