Skip to content
MeghaOS

Select language

MeghaOS speaks over 100 languages on your own machine. This site is available in full in every language listed here; our legal pages and blog posts stay in English.

FIG 8.0 · Model card

Small enough to runwhere your data already is.

Megha is a 0.6B-parameter causal language model built for agentic workflows, coding tasks and multilingual instruction-following. The size is the point: local processing is only a real guarantee if the model actually fits on your machine.

FIG 8.1 · Specification

The whole model card, on one screen.

Parameters
0.6B
total
Non-embedding
0.44B
parameters
Layers
28
transformer blocks
Attention heads
16Q / 8KV
grouped-query
Context length
32,768
tokens
Architecture
Causal LM
decoder-only
Languages
100+
instruction-following
Thinking mode
Toggleable
at inference time
FIG 8.2 · Schematic

28 layers, grouped-query attention.

Grouped-query attention, with 16 query heads sharing 8 key/value heads, is what keeps the KV cache small enough for a 32K context on consumer hardware.

01
02
03
04
05
06
07
08
09
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
embedding · 0.16B
Per block

Grouped-query attention · 16Q / 8KV

16 query heads8 key/value heads · shared

Sharing key/value heads across query heads cuts the KV cache roughly in half against full multi-head attention. That is what makes a 32K context viable on a laptop rather than a workstation, and the reason the local-processing guarantee is practical instead of theoretical.

FIG 8.3 · Thinking mode

Reasoning you can switch on when it earns its cost.

Thinking mode is toggleable at inference time rather than baked into the model. Routine extraction, classification and formatting run without it and stay fast; multi-step planning and difficult refactors get the extra reasoning budget. The agent scheduler makes this choice per task, and you can override it.

  • Toggled per request, not per deployment
  • Routine work stays low-latency by default
  • The agent scheduler decides, and shows you what it decided
  • Budgets are enforced either way; thinking mode is not a blank cheque
When it engages
  • OffClassification, extraction, formatting, template completion
  • OffSimple routing and embedding generation
  • OnMulti-file refactors and dependency-aware edits
  • OnMulti-step planning where a wrong first step is expensive
FIG 8.4 · What actually runs

Megha is a choice, not a default you cannot change.

MeghaOS bundles a small model so the first launch works offline on modest hardware. Megha, and anything else you prefer, is an install away.

A compact model is bundled

So the first launch works offline on any supported machine, including older Macs and PCs with limited memory. It handles routing, extraction, classification and tool selection well, and it is deliberately small rather than deliberately capable.

Megha is a one-line install

Pull it through Ollama or serve it with vLLM and point MeghaOS at the endpoint. This is the configuration the model card above describes, and it is what we recommend on any machine with memory to spare.

Or bring a different model entirely

Anything you can serve through Ollama or vLLM works: Qwen, DeepSeek, Kimi, or whatever you already run. MeghaOS talks to an OpenAI-compatible endpoint and does not care what is behind it.

Or use a hosted provider

Point it at a cloud API with your own key, or subscribe to MeghaOS Pro. Requests you route this way leave the device, which is the whole trade. See pricing for what that costs and what it buys.

FIG 8.5 · Deployment

Run it outside MeghaOS if you want to.

The model integrates with the runtimes you already use. Nothing about it is locked to the operating system.

Ollama
ollama run megha
vLLM
vllm serve meghaos/megha-0.6b
SGLang
python -m sglang.launch_server --model-path meghaos/megha-0.6b
Transformers
from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained("meghaos/megha-0.6b")

Model weights and tokenizer configs available at hf.co/meghaos/megha-0.6b under Apache 2.0.

FIG 8.6 · Limits

What a 0.6B model is not.

Being straight about this is the reason to trust the rest of the page.

It is not a frontier model

On open-ended reasoning, long-form writing and hard novel problems, a 0.6B model will not match a frontier system. It is not trying to.

The trade is deliberate

For classification, extraction, routing, formatting and tool-calling, which is most of what an agent actually does, a small local model is faster, free per call, and immune to rate limits.

Hybrid is supported, not required

You can point MeghaOS at a larger local model or a hosted one. Anything you route to a hosted provider leaves the device. That is what routing there means, and it is a setting you choose rather than something that happens quietly.

Local removes a whole failure class

A large share of production LLM errors are rate limits. A local model cannot be rate limited, so an agent can keep working during a window when a hosted one would stall.

Run the model wherethe data already lives.

MeghaOS runs a compact model out of the box and lets you swap in Megha, another open model, or a hosted provider whenever you want.