New to AI agents or speech tech? Start here (1 min)
I already had an AI assistant I type to for reminders, quick questions, and a daily summary. The problem was the car, where I have no screen or keyboard to spare.
So I built a phone app that listens when I tap, turns my speech into text, sends it to the assistant, and reads the answer aloud. Everything except the assistant's language model runs on a small computer in my house. My voice never leaves my home network.
The first version worked at a desk and failed in the car. The phone's mute switch silenced the replies, and when the phone lost signal the app went quiet with no message. Both got fixed.
The hard part was teaching it to know when I had finished talking over road noise. A tiny speech-detection model solved that, after one bug that returned wrong answers without ever raising an error.
The lesson: the unglamorous details (right audio channel, clear failure messages, a spoken way to say goodbye) mattered more than the AI itself.
I speak
→ speech becomes text
→ the assistant thinks
→ text becomes speech
→ I hear the replyGlossary at the end.
I already had an AI agent I talk to for day-to-day things: reminders, quick lookups, a daily briefing. Call it Hermes. The problem was that it only understood text on a screen, and the place I most wanted to talk to it was in the car, where I don't have a screen or a keyboard to spare.
So I built a voice front end for it: a phone-installable web app that listens, sends what I said to the agent, and speaks the reply back. Everything runs on a small headless server I already had on my home network. No cloud speech-to-text (STT), no cloud text-to-speech (TTS), nothing paid beyond the DeepSeek API the agent already used. This is what that build looked like, including the parts that didn't work the first time.
The shape of it
Four stages, all local except the actual "thinking" step:
phone mic
→ STT
→ the agent (→ language model)
→ TTS
→ phone speaker- STT:
faster-whisper, running on CPU. A short question transcribes in a couple of seconds. - The agent: the existing text-based agent I already had. The voice app reaches it through a one-shot CLI bridge, shelling out to the agent's command-line interface once per turn, with a
--continueflag that keeps a single ongoing conversation across turns. - TTS: two local engines behind a toggle, Kokoro for a natural voice and Piper for a faster, lighter one, so I can A/B them by ear.
- The client: a React Progressive Web App (PWA) built with Vite, talking to a FastAPI server; tap-to-talk, added to the home screen so it behaves like a real app instead of a browser tab.
Access from the phone goes over Tailscale back to the home server, so nothing is exposed to the open internet and the only outbound call anywhere is the agent's own model API. Audio never leaves my own network.
Getting it to work in a moving car
The first version worked fine sitting at a desk and fell apart the moment I drove with it. Two problems showed up immediately:
The silent switch muted replies. iOS routes WebAudio playback through the ringer channel by default. So flipping the physical mute switch, which I do reflexively getting into the car, silenced Hermes along with everything else. The fix was switching playback to a plain <audio> element, which plays on the media channel like a music app and ignores the ringer switch entirely.
Cellular failures looked like nothing. On Tailscale, if the phone's tunnel drops, which happens more than you'd think on cellular, the app just sat there. No error, no feedback, just silence where a reply should be. I added a lightweight health check that polls the server and shows a clear "can't reach the server" banner instead of failing invisibly. Small fix, but it was the single biggest trust-builder. A tool that fails loudly is one you keep using; one that fails silently is one you stop trusting after the second time.
I also added a hands-free mode: tap once, and the mic re-arms itself after each reply, so a whole conversation is one tap instead of one tap per turn. A spoken exit ("goodbye Hermes") ends the session without my looking at the screen.
Replacing the ears
For weeks the weak point was turn-taking: deciding when I'd finished talking. That ran entirely in the browser, measuring the microphone's volume and treating a stretch of quiet as "done." At a desk it worked. In a car it did not. Road noise sits in the same volume range as speech, so the threshold was a constant compromise between cutting off soft words and refusing to stop on highway hum.
Then I found an open-source project doing something adjacent, a full local speech-to-speech pipeline, and asked whether to just switch to it. The answer was no. It is a desktop app that owns the machine's own microphone and speaker, so it does not solve my problem: reaching the agent from my phone, hands-free, over a network. Its brain is a generic chat endpoint with no memory, no session continuity, no tools. Adopting it would have thrown away the parts of Hermes that make it useful.
One piece of it was a real upgrade, though. Instead of a volume threshold, it does turn-taking with a small neural model for voice activity detection (VAD): Silero VAD, running via ONNX Runtime on CPU, trained to tell speech apart from everything else. That is exactly what my hand-tuned threshold was standing in for, badly. So I lifted just that.
Adding server-side turn detection, without breaking what worked
The rule I set for myself: the existing tap-to-talk and hands-free modes had to keep working exactly as they were.
So instead of replacing anything, I added a second, parallel way of talking to it: a "continuous" mode, opt-in, off by default. The original mode still uploads one recording per utterance and decides "you're done talking" with a volume timer in the browser, unchanged. The new mode opens a WebSocket, streams raw audio continuously, and lets the server decide when I'm done talking, using the VAD model instead of a threshold. Both paths funnel into the same downstream pipeline, so they can't quietly drift apart over time.
The VAD model itself is tiny (a few megabytes) and runs fast enough on CPU that it's not the bottleneck; transcription and the model call still dominate response time by a wide margin.
The bug that taught me the most
Getting the new model wired in, I ran it against a clean, loud, unambiguous recording of speech and got back a probability near zero, over and over. Not "the audio is noisy so it's uncertain," but confidently wrong, on audio that should have been trivial.
The model's input shape was declared as flexible: it would accept a chunk of any length without complaint. I was feeding it exactly the window size the documentation described. That turned out to be the trap. The model expects each chunk plus a small sliver of context carried over from the previous chunk, a few dozen samples' worth of continuity between calls. Feed it the window alone, with no error and no warning, and it silently returns garbage forever. The interface was permissive enough to hide a mistake that would otherwise have been a one-line fix.
I only caught it because I refused to trust "it imported fine and didn't crash" as a correctness signal. Instead I ran a real recording through it and looked at the actual numbers coming out. It's the same lesson as the silent cellular-failure bug from months earlier, wearing different clothes: the failure mode that costs you the most time is never the one that throws an exception. It's the one that returns a plausible-looking value and lets you walk away thinking it worked.
Once that fix was in, the difference was obvious. Feed it road noise alone, it correctly says "not speech." Feed it speech mixed with road noise, it correctly still says "speech." Feed it two sentences with a short pause between them, it correctly waits for the real pause instead of splitting the first one off mid-thought. None of that was reliable with a volume threshold.
What I'd tell someone doing this themselves
Build the interface first, the intelligence second. The plumbing that made this usable day to day mattered more than which speech model was underneath.
Adding a capability isn't the same as replacing one. When something already works, run the new thing alongside the old until it has earned the swap, and design them so the two can't quietly diverge.
A component that loads without error has told you nothing about whether it works. The most expensive bug here didn't crash; it returned a wrong number and nothing objected. Find that class of bug by feeding the system an input whose answer you already know, and checking.
Hermes drives with me now: a small voice loop between a phone and a box in a closet, and a reminder that the unglamorous 80% is where the product usually lives.
Glossary
- AI agent: an AI that takes multi-step actions with tools rather than only answering questions; here, the assistant the voice app talks to.
- speech-to-text (STT): turning recorded speech into written text, done here on the home server by faster-whisper.
- text-to-speech (TTS): turning the assistant's written reply into spoken audio, done here by Kokoro or Piper.
- voice activity detection (VAD): a model that decides whether a slice of audio contains speech, used here to tell when I have stopped talking.
- noise floor: the background sound level when nobody is speaking, which the old volume-based approach measured to decide what counted as quiet.
- PWA: a Progressive Web App, a website that can be added to a phone's home screen and behaves like an installed app.
- WebSocket: a connection between the phone and the server that stays open, so audio can stream continuously instead of being uploaded in separate requests.
- mesh VPN: a private network that links my devices directly to each other over the internet (Tailscale here), so the phone reaches the home server without exposing it publicly.
- false positive: when the system decides speech is present, or a turn has ended, when it has not.
- model / inference: a model is the trained neural network; inference is running it on new input to get an answer, such as a speech probability.
- latency: the delay between finishing a sentence and hearing the reply, dominated here by transcription and the language-model call.