Skip to content
MeghaOS

Select language

MeghaOS speaks over 100 languages on your own machine. This site is available in full in every language listed here; our legal pages and blog posts stay in English.

Guide· MeghaOS· 6 min read

Running AI agents locally: a practical guide

What hardware you need, how quantisation and context length drive memory, which model sizes suit which work, and where local inference stops being the right answer.

Running a language model on your own machine went from a research exercise to a reasonable default faster than most people noticed. The tooling is good, the small models are genuinely capable, and consumer hardware from the last few years is enough for real work.

This guide covers what determines whether a model fits, what performance to expect, which sizes suit which tasks, and, the part usually left out, when local inference is the wrong choice.

What determines whether a model fits

Almost everything reduces to memory. Three things consume it.

Weights. The parameters, at whatever precision you load them. A 7-billion-parameter model at 16-bit precision needs roughly 14 GB just for weights. The same model quantised to 4-bit needs roughly 4 GB. This is the dominant term and the one quantisation addresses.

KV cache. Every token in the context window has cached key and value tensors, and this scales with context length. It is the term people forget. A model that loads comfortably can still run out of memory at long context, because the cache grows as the conversation does. Agentic workloads hit this hard: tool results and file contents fill a window quickly.

Overhead. The runtime, the compute buffers, and whatever else the machine is doing. Budget a couple of gigabytes.

A workable rule: take the parameter count in billions, multiply by the bytes per parameter implied by your quantisation, add 20–30% for cache and overhead at moderate context, and require that to fit in available memory with room to spare. Swapping to disk does not degrade inference gracefully; it stops it being usable.

Quantisation, briefly

Quantisation stores weights at lower precision. It is the single highest-leverage knob.

  • 8-bit is close to indistinguishable from full precision for most purposes, at roughly half the memory.
  • 4-bit is where most local deployment lands. Quality loss is measurable but modest, and memory drops to roughly a quarter.
  • Below 4-bit degrades noticeably, and how much depends on the model. Worth testing rather than assuming.

The practical implication is counterintuitive and worth internalising: a larger model at 4-bit usually beats a smaller model at 8-bit for the same memory budget. Parameters bought with precision are generally a good trade.

Hardware

Apple Silicon is unusually well suited to this because memory is unified: the GPU addresses the same pool as the CPU, so a 32 GB Mac can put far more into a model than a discrete GPU with 8 GB of VRAM. Memory bandwidth varies substantially across the Pro, Max and Ultra tiers, and bandwidth is what sets generation speed once the model fits.

Discrete NVIDIA GPUs are fastest when the model fits in VRAM, and fall off a cliff when it does not. VRAM capacity matters more than raw compute for this workload. A card with less memory and more compute will lose to one with more memory and less.

CPU-only works and is slower, usually by a large factor. For a small model doing short completions it is tolerable. For agentic work involving many sequential model calls, it usually is not.

The honest framing: memory capacity determines what you can run at all, memory bandwidth determines how fast it runs, and compute is rarely the binding constraint for single-user inference.

Matching model size to work

Bigger is not automatically better, because latency compounds. An agent making twenty sequential calls turns a two-second-per-call model into a forty-second task.

Sub-1B models are for classification, routing, extraction, short structured outputs, and deciding which tool to call. They run fast on almost anything and are frequently the right component inside a larger system rather than the whole system. Megha 0.6B is built for this role: agentic workflows, tool selection, and multilingual instruction-following, small enough to coexist with the rest of your machine’s work. MeghaOS bundles something smaller still so a first launch works offline on older hardware, and lets you point it at Megha, or any model you serve through Ollama or vLLM, once you know what the machine can carry.

3B–8B models handle general assistant work, summarisation, and straightforward code. This is the sweet spot for a laptop, and where most people should start.

13B–34B models are noticeably better at multi-step reasoning and harder code, and need serious memory: realistically a well-specified desktop or a high-memory Mac.

70B and above approach hosted-model quality and are impractical on typical consumer hardware at usable speed.

The design that works best in practice is not one model. It is a small fast model handling routing, extraction and the many small decisions an agent makes, escalating to something larger only for the steps that need it. Most of an agent’s model calls are not the hard part.

What performance feels like

Two numbers matter and they are not interchangeable.

Time to first token is how long before output starts. This governs whether the system feels responsive, and it is dominated by prompt processing, so it grows with context length. An agent that stuffs a large file into context pays this on every call.

Tokens per second is the generation rate. Above about 20 it reads faster than most people; between 10 and 20 is usable; below 5 is uncomfortable for interactive work though fine for background jobs.

For agentic work, time to first token usually matters more than people expect, because agents make many calls with large contexts and the user waits through all of them.

Where local inference is the wrong answer

Being straightforward about this is more useful than advocacy.

Genuinely hard reasoning. A frontier model is better at difficult multi-step problems than anything you will run on a laptop, and the gap is real. Where the task is hard rather than context-heavy, hosted wins.

Long-context work at speed. Prompt processing over very large contexts is slow locally. Hosted infrastructure has hardware you do not.

Bursty multi-user load. Local inference serves one machine. Shared load is what servers are for.

Machines doing other work. A model resident in memory is memory unavailable to everything else. On a laptop already running a browser and an IDE, this is a real cost.

The reasonable architecture keeps both available and makes the choice explicit. Local by default, because that is where your context is and it costs nothing per request. Hosted for the specific steps that need it, with the consequence stated: that request leaves the device. How MeghaOS splits this is the same distinction: the operating system and local model are free because they run on your hardware; the only thing charged for is inference someone has to pay for.

Getting started

The shortest useful path:

  1. Work out your memory budget. Total RAM, minus what your machine needs to keep running, minus a couple of gigabytes of headroom.
  2. Pick the largest 4-bit model that fits with room for the KV cache at the context length you actually use.
  3. Measure both numbers, time to first token and tokens per second, on a prompt resembling your real work, not a one-line test.
  4. If it is too slow, reduce context before reducing model size. The cache is often the problem.
  5. If quality is short, try a larger model at lower precision before a smaller one at higher precision.

For agentic work specifically, instrument how many model calls a task makes before optimising the calls themselves. The usual finding is that most calls are small decisions that a much smaller model handles at a fraction of the latency, and that restructuring the agent beats upgrading the model.


Related: The Megha 0.6B model card · What a local model changes day to day · Local-first is a capability argument

Written by

MeghaOS, building a Wayland-native operating system designed to host agentic AI on hardware you own. More about us.

Run it on your own machine.

Free to download. Nothing leaves the device unless you connect it. Enterprise deployment is a conversation away.