Small enough to runwhere your data already is.
Megha is a 0.6B-parameter causal language model built for agentic workflows, coding tasks and multilingual instruction-following. The size is the point: local processing is only a real guarantee if the model actually fits on your machine.
The whole model card, on one screen.
- Parameters
- 0.6B
- total
- Non-embedding
- 0.44B
- parameters
- Layers
- 28
- transformer blocks
- Attention heads
- 16Q / 8KV
- grouped-query
- Context length
- 32,768
- tokens
- Architecture
- Causal LM
- decoder-only
- Languages
- 100+
- instruction-following
- Thinking mode
- Toggleable
- at inference time
28 layers, grouped-query attention.
Grouped-query attention, with 16 query heads sharing 8 key/value heads, is what keeps the KV cache small enough for a 32K context on consumer hardware.
Grouped-query attention · 16Q / 8KV
Sharing key/value heads across query heads cuts the KV cache roughly in half against full multi-head attention. That is what makes a 32K context viable on a laptop rather than a workstation, and the reason the local-processing guarantee is practical instead of theoretical.
Reasoning you can switch on when it earns its cost.
Thinking mode is toggleable at inference time rather than baked into the model. Routine extraction, classification and formatting run without it and stay fast; multi-step planning and difficult refactors get the extra reasoning budget. The agent scheduler makes this choice per task, and you can override it.
- Toggled per request, not per deployment
- Routine work stays low-latency by default
- The agent scheduler decides, and shows you what it decided
- Budgets are enforced either way; thinking mode is not a blank cheque
- OffClassification, extraction, formatting, template completion
- OffSimple routing and embedding generation
- OnMulti-file refactors and dependency-aware edits
- OnMulti-step planning where a wrong first step is expensive
Megha is a choice, not a default you cannot change.
MeghaOS bundles a small model so the first launch works offline on modest hardware. Megha, and anything else you prefer, is an install away.
A compact model is bundled
So the first launch works offline on any supported machine, including older Macs and PCs with limited memory. It handles routing, extraction, classification and tool selection well, and it is deliberately small rather than deliberately capable.
Megha is a one-line install
Pull it through Ollama or serve it with vLLM and point MeghaOS at the endpoint. This is the configuration the model card above describes, and it is what we recommend on any machine with memory to spare.
Or bring a different model entirely
Anything you can serve through Ollama or vLLM works: Qwen, DeepSeek, Kimi, or whatever you already run. MeghaOS talks to an OpenAI-compatible endpoint and does not care what is behind it.
Or use a hosted provider
Point it at a cloud API with your own key, or subscribe to MeghaOS Pro. Requests you route this way leave the device, which is the whole trade. See pricing for what that costs and what it buys.
Run it outside MeghaOS if you want to.
The model integrates with the runtimes you already use. Nothing about it is locked to the operating system.
ollama run meghavllm serve meghaos/megha-0.6bpython -m sglang.launch_server --model-path meghaos/megha-0.6bfrom transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("meghaos/megha-0.6b")Model weights and tokenizer configs available at hf.co/meghaos/megha-0.6b under Apache 2.0.
What a 0.6B model is not.
Being straight about this is the reason to trust the rest of the page.
It is not a frontier model
On open-ended reasoning, long-form writing and hard novel problems, a 0.6B model will not match a frontier system. It is not trying to.
The trade is deliberate
For classification, extraction, routing, formatting and tool-calling, which is most of what an agent actually does, a small local model is faster, free per call, and immune to rate limits.
Hybrid is supported, not required
You can point MeghaOS at a larger local model or a hosted one. Anything you route to a hosted provider leaves the device. That is what routing there means, and it is a setting you choose rather than something that happens quietly.
Local removes a whole failure class
A large share of production LLM errors are rate limits. A local model cannot be rate limited, so an agent can keep working during a window when a hosted one would stall.
- Where the model sits in the architectureThe six layers between the hardware and the composed interface.
- Local inference vs. hosted inferenceWhat is free because it runs on your hardware, and what is not.
- What a local model changes day to dayThe capability argument for running the model where the data is.
Run the model wherethe data already lives.
MeghaOS runs a compact model out of the box and lets you swap in Megha, another open model, or a hosted provider whenever you want.