Skip to content
MeghaOS

Select language

MeghaOS speaks over 100 languages on your own machine. This site is available in full in every language listed here; our legal pages and blog posts stay in English.

FIG 3.9 · Voice

Talk to it.Nothing leavesthe room.

Voice assistants have always come with a microphone pointed at someone else's servers. This one does not. Speech recognition and speech synthesis both run on your machine, which makes voice usable for the conversations you would never have near a smart speaker.

The reason you do not use voice is the reason this exists.

Every mainstream voice assistant records you, sends the audio to a data centre, and keeps it for some period described in a document you did not read. So people use voice for timers and weather, and type everything that matters.

Move recognition onto the machine and that calculation disappears. What is left is the fastest input method humans have, pointed at an agent that can actually act. We wrote up the longer argument in on-device speech recognition, and why voice assistants stayed useless.

What never happens
  • Audio uploaded to a recognition service
  • A transcript stored on someone else’s disk
  • A per-minute transcription bill
  • Voice stopping because the connection dropped
  • An account required to speak to your own computer
FIG 3.10 · End to end

Four steps, none of which leave the machine.

Speech in, speech out, with the same agent and the same tools in between.

01

It notices you started talking

Speech is separated from room noise on the device, so the microphone is not streaming anything anywhere while you are not addressing it.

02

Your words become text, here

Recognition runs on your own processor. No audio is uploaded, and there is no transcription service holding a copy of what you said.

03

The agent does the work

The same agent, the same tools, the same composed interfaces. Voice is another way in, not a reduced version of the product.

04

It answers out loud

Speech synthesis runs locally too, so a reply starts without a round trip and keeps working with the network unplugged.

FIG 3.10b · Spoken in, composed out

You asked out loud. This is what you get back.

Not a sentence read at you, but the same composed interface a typed request would produce, with the reading-aloud offered rather than assumed. Drawn by the desktop's own renderer; only the messages are invented.

Read me what came in overnight

try

Overnight, on one screen

You asked out loud. The answer is a screen, not a paragraph read back at you.

Came in
31messages
Actually for you
4
↑ the rest is noise

The build went red at 02:14 and green again at 02:51; someone had already fixed it before you woke up.

One contract came back signed. Two threads are waiting on a number only you have.

Wants you

Pricing thread
Waiting since 23:40
2
Contract, signed
Nothing to do
✓
Build recovered
Fixed at 02:51
✓

Or have it read back

Speech synthesis runs here too, so nothing about this leaves the machine either.

Voice
Pace105%
50200
Keep listening after it finishes
FIG 3.11 · Not a lesser mode

Voice reaches the whole system, not a cut-down version of it.

Most assistants give voice a smaller brain than the app: a limited command set, a handful of intents, and a shrug for anything else. Here it is the same agent with the same tools, so a spoken request can drive the browser, edit files, or compose an interface exactly as a typed one would.

  • The same agent, the same tools, the same memory
  • A spoken request can return a composed interface, not just a sentence
  • Long jobs keep running and tell you when they are done
  • Switch between speaking and typing mid-task without losing the thread
Where it fits
  • Hands busy, eyes elsewhere
    Ask for the state of something while you are doing something else, and get a spoken answer plus a composed view waiting when you look back.
  • Dictation that understands context
    Not just transcription: the agent has your files and your history, so "write that up for the Tuesday thread" resolves to the right thing.
  • Genuinely private conversations
    The kind of thing people will not say to a smart speaker. Legal, medical, financial, personal. Nothing is uploaded, so nothing is retained by anyone.
  • Works with no connection
    On a plane, in a facility with no network, on a machine that is deliberately air-gapped. Voice does not degrade to unavailable.
FIG 3.12 · Questions

What people ask about voice.

Does my voice get sent to a server?

No. Recognition and synthesis both run on your machine, using models that ship with the system. There is no audio upload step and no transcription provider in the path, which is the whole reason voice is usable for the conversations people otherwise avoid having near a microphone.

Does voice work offline?

Yes. The speech models are installed locally alongside the language model, so a machine with no network connection still listens and answers. This is the same property that makes an air-gapped install viable.

Do I need an account or a subscription for voice?

No. Voice is part of the local system and costs nothing per use. There is no per-minute transcription bill because there is no transcription service.

Is it always listening?

Speech is separated from background noise on the device, and nothing leaves the machine at any point, so even while it is listening for you, there is no stream going anywhere to be intercepted or retained.

Can I use voice and the screen together?

That is the intended way to use it. A spoken question can return a spoken answer and a composed interface at the same time, so you get the short version out loud and the detail on screen when you turn back to it.

Say it out loud.

Free to download, and the microphone stays pointed at your own machine.