Skip to content
MeghaOS

Select language

MeghaOS speaks over 100 languages on your own machine. This site is available in full in every language listed here; our legal pages and blog posts stay in English.

Product· MeghaOS· 6 min read

On-device speech recognition, and why voice assistants stayed useless

Voice is the fastest input humans have, and we use it for timers. The reason is not accuracy. It is that every mainstream assistant points a microphone at somebody else's data centre.

Speech is roughly three times faster than typing, needs no surface, and works while your hands are doing something else. By any reasonable measure it should be the default way to talk to a computer.

Instead, a decade after voice assistants shipped to everyone, the median use is setting a timer. Not because recognition is bad, since it has been good for years, but because of where the recognition happens.

The problem is not accuracy

The standard explanation for why voice failed is that it misunderstands people. This was true and has largely stopped being true. Modern recognition handles accents, crosstalk and domain vocabulary well enough that transcription errors are no longer the thing standing between you and using it.

Watch what people actually do instead. They use voice for requests where being overheard costs nothing (timers, weather, music, navigation) and they type everything else. That is not the behaviour of someone who does not trust the transcription. It is the behaviour of someone who does not trust the destination.

This is the same withholding that limits cloud assistants generally, sharpened by the medium. Typing into a text box at least feels deliberate. A microphone in the room is ambient, and the question “what happens to this audio” applies continuously rather than per request.

So the assistant is confined to the class of request where context is worthless, and it is exactly as useful as that confinement implies.

What “on-device” has to mean

It is worth being precise, because the phrase gets applied loosely.

A voice interface has three moving parts. Something decides you are speaking rather than the room being noisy. Something turns speech into text. Something turns text back into speech.

Vendors routinely put the first part on the device, since the wake-word detector runs locally, which is what makes “it only listens after the wake word” technically true, and then send everything after it upstream. That is the architecture most people are actually using. The local component is a gate, not a boundary. Once it opens, audio leaves.

On-device in the sense that matters means all three run on your processor, and no audio is transmitted at any point. Not buffered pending upload, not sent for quality review, not retained for model improvement. The distinction is not a policy difference. It is whether the network path exists.

Why this became possible recently

Running recognition locally used to be a genuine trade. On-device models were noticeably worse than server-side ones, and the gap was large enough that the privacy argument lost to the accuracy argument for most people.

Two things closed it. Model architectures got dramatically more efficient at small sizes, so a recognition model that fits comfortably in a fraction of a laptop’s memory now performs close to what required a server a few years ago. And ordinary consumer hardware acquired enough parallel compute that running one is not the machine’s main activity.

Speech synthesis followed the same curve, and further, because good synthetic speech is a much smaller problem than good recognition, and local voices crossed the line from robotic to unremarkable before recognition did.

The result is that “local speech is worse” is now a claim about degree rather than kind, and the degree is small enough that it stops driving the decision.

What changes when the audio does not leave

Latency stops being round-trip bound. A hosted assistant cannot respond faster than the network allows, and the pause before a reply is the single thing that most makes a voice interface feel like a machine. Local recognition removes the round trip entirely, which changes the texture of the interaction more than any accuracy improvement would.

It works with no connection. On a plane, in a basement, in a facility that does not permit outbound traffic, on a machine deliberately kept off the network. Voice stops being a feature that degrades to unavailable at the moment you are least able to type.

There is no per-minute bill. Hosted transcription is metered, which is why products built on it ration it: short sessions, a cap on dictation length, a paid tier for the thing you wanted. Local recognition costs electricity, so the design question becomes what is useful rather than what is affordable.

The wake-word theatre becomes unnecessary. Most of the ceremony around voice assistants exists to reassure you about upload boundaries. Remove the upload and the reassurance has nothing to do.

The set of things you will say out loud expands. This is the one that matters. The requests worth automating are the ones involving your contracts, your finances, your health, your work in progress: precisely the ones nobody says near a smart speaker.

Voice alone is a bad interface

The correct criticism of voice is not privacy. It is bandwidth in the other direction.

Speaking is fast. Listening is slow, linear, and impossible to skim. A spoken answer cannot show you a table, cannot let your eye jump to the row you care about, and cannot be re-read without replaying the whole thing. Any question whose honest answer is structured gets flattened into a paragraph, or worse, into “I found several results.”

This is why voice-only devices plateaued at commands rather than progressing to work. The input channel was good enough for real requests and the output channel was not.

The fix is not better speech. It is to stop treating voice as a complete interface and treat it as the fastest available input to a system that can answer on a screen. Ask out loud, get a one-sentence spoken acknowledgement, and find the actual answer composed as an interface you can look at when you turn back. Speech for the request, structure for the response.

That requires the assistant to be able to generate an interface rather than only text, which is a separate capability from speech and the reason most voice products cannot do it.

What this does not fix

Local recognition is not magic, and the honest limits are worth stating.

Accented speech, heavy background noise and specialist vocabulary are still harder than clean speech in a quiet room, and a small local model has less headroom for them than a large hosted one. The gap is narrow now; it is not zero.

Speech is still a poor fit for anything requiring precision: exact identifiers, code, anything you would need to spell. Dictating a UUID is not a solved problem and will not become one.

And running locally says nothing about what happens after transcription. If the resulting text is sent to a hosted model, the audio stayed home and the content did not. The guarantee is only as strong as the weakest link in the whole path, which is an argument for running the model locally too rather than an argument against local speech.

Where MeghaOS sits

Voice in MeghaOS runs entirely on the machine: activity detection, recognition and synthesis, with no audio upload step in the path. It reaches the same agent as the text interface, with the same tools and the same memory, so a spoken request can drive the browser or edit files rather than selecting from a command list.

And because the system composes interfaces rather than emitting text, a spoken question can return a spoken summary and a structured view at the same time, which is the arrangement the medium actually wants.

The result is a voice interface you can use for the things you would otherwise type, which, given how much faster speaking is, is the entire point.


Related: Voice, end to end · Why local-first is a capability argument · Running AI agents locally

Written by

MeghaOS, building a Wayland-native operating system designed to host agentic AI on hardware you own. More about us.

Run it on your own machine.

Free to download. Nothing leaves the device unless you connect it. Enterprise deployment is a conversation away.