Magnitude or llama.cpp? Which runs faster on your machine based on your hardware

Magnitude runs local models up to 2× faster than llama.cpp on some machines and slower on others: who wins depends on your chip, not the headline. Magnitude is an open source inference engine from Magnitude AI Inc. (YC S25), launched on Hacker News on September 30. It profiles your hardware, recommends models that fit, tunes their GPU kernels on your own machine, and connects to the agents you already use: Claude Code, OpenCode, Codex, Cline, Hermes, or OpenClaw.

The manufacturer’s numbers are real, published, and reproducible. The independent numbers that appeared in the first 24 hours tell a more interesting story. If you run models locally, read both before you switch.

What is Magnitude?

Magnitude is a local inference server designed for agent workloads: long sessions, multiple agents at once, and a machine you keep using for other things in the meantime. It’s written in Rust, with its own GPU kernel runtime and autotuner. It’s licensed under Apache 2.0 and, as of October 1, 2026, the repository had 6.1k stars and 409 forks.

It’s distributed as a desktop app for macOS, Windows, and Linux. You choose a model in the app, press Connect next to your agent, and Magnitude loads the model when the agent requests it and unloads it when it goes idle or memory fills up.

The project has history: the same team previously built an open source browser agent, and the repository evolved from that product to inference. According to the founders, they created Magnitude because no existing engine worked well for running agents locally.

How does it know what model your machine can run?

Magnitude profiles your chip, memory, and memory bandwidth, and estimates the fit and tokens per second of each model in the catalog before you download anything. Then it orders them by speed, accuracy, intelligence, and memory. In the app’s Discover view you start from a “Balanced” recommendation and slide a control toward speed or toward intelligence.

Two clarifications from the maintainers themselves on Hacker News:

  • Speed estimates are estimates. They’re useful for filtering and sorting models, they don’t include the gain from speculative decoding, and some users saw figures well below what they measure in practice. The team says it’s working to make this clearer.
  • Tuning is what changes real performance. It’s done once for each model you download and, according to the founders, takes about a minute.

The catalog only includes models quantized to 4 bits or higher. Below that, according to the team, reasoning, tool calls, and coherence degrade too much to work with agents.

Is Magnitude really 2× faster than llama.cpp?

In the manufacturer’s own benchmark, decode speed nearly doubles on an M4 Pro Mac. On other hardware the advantage varies, and some users measured it slower than llama.cpp.

These are the figures Magnitude published at launch. All are self-reported:

Machine Decode Prefill Memory per agent
Mac M4 Pro 48 GB (Metal) 30 → 57 tok/s (+92%) 466 → 507 tok/s (+9%) 28% less
DGX Spark (CUDA) 49 → 58 tok/s (+19%) 2,033 → 2,507 tok/s (+23%) 27% less

The test used Qwen 3.6 35B A3B at 4 bits, with 64k context and no speculative decoding. According to the founders, the task puts Moby Dick in the context and asks the model to repeat the last section. The benchmark code is public in the repository.

There’s a methodological detail that matters. Magnitude quantizes its KV cache to 8 bits for keys and 4 bits for values. The team says they tried the same configuration in llama.cpp and decode tanked, so llama.cpp ran with 16-bit KV cache. Part of the advantage comes, therefore, from that quantization and not just from the kernels. The team claims their RULER-based retrieval tests show no quality loss on long context; that claim is also theirs.

On the first day, several Hacker News users published their own measurements:

  • RTX 5070 Ti: llama.cpp was 20% to 30% faster on decode.
  • M5 Max: llama.cpp was about 2× faster on prefill and decode with Qwen 3.8.
  • M5 Max against rapid-mlx (MLX): 175 tok/s decode versus 161 for Magnitude, and MLX delivered the first token sooner.
  • M5 Pro against oMLX: Magnitude was faster on decode (82.8 versus 76.5 tok/s), but about 2.6× slower on prefill.

The maintainers acknowledged a known issue on M5 chips: their kernels still don’t take advantage of Metal 4’s matrix multiplication hardware. As of October 1, 2026 they described the patch as “in progress.”

The practical takeaway: Magnitude’s advantage is clearer where it was tuned and measured, namely on M4-class Apple Silicon. On an M5 or some NVIDIA cards, measure before you switch.

Magnitude, Ollama, LM Studio, or MLX?

Magnitude is the option that picks the model and tunes kernels for you. Ollama and LM Studio prioritize broad compatibility, and MLX is the specialized path for Apple Silicon.

The team itself groups existing engines into three families:

  • Batch servers designed for data centers: vLLM, SGLang.
  • Broad compatibility engines, which trade maximum speed for coverage: llama.cpp, Ollama.
  • Specialized engines, fast on one platform but less complete: oMLX and similar.

Magnitude tries to tune generic kernels to your specific device.

On Mac, several Hacker News commenters noted that the fair comparison is against MLX, not llama.cpp. The team says it will publish benchmarks against MLX; as of October 1, 2026 it hadn’t yet.

If you’re choosing your local stack from scratch, start with our comparison LM Studio, Ollama, or llama.cpp: what really runs on your machine. Magnitude is the new candidate in that same decision.

Does it work on Windows, Mac, and Linux?

Yes, on all three. The documentation lists these configurations:

Device Inference
Mac with Apple Silicon GPU with Metal
Mac with Intel CPU only
Windows x64 (10 version 1809 or later, or 11) CPU, NVIDIA CUDA, or Vulkan
Linux x64 / ARM64 (glibc 2.35 or later) CPU, NVIDIA CUDA, or Vulkan

Four limits worth knowing:

  • NVIDIA: CUDA builds target Ampere GPUs or newer; older cards don’t have automatic support.
  • AMD: AMD GPUs work with Vulkan; there’s no ROCm backend.
  • Multiple GPUs: as of October 1, 2026, they’re not supported; the team has it on their near-term roadmap.
  • Linux: RHEL 9 is out because its glibc is 2.34, and there are no packages for Alpine/musl.

The GitHub README still describes Windows as “via WSL only”. Current documentation describes a native Windows installer. Follow the documentation.

Is Magnitude free?

Yes. The engine is Apache 2.0, runs entirely on your machine, and has no token costs, API keys, or usage limits. Once you download the model, it works offline.

On Hacker News, the founders described their business model as a future paid inference cloud for hybrid workloads. Local models would do what they can, and a token-charged cloud would solve the harder tasks, switching between them without losing the prefix cache. That cloud doesn’t exist yet.

How do you install Magnitude?

Download the desktop app from magnitude.dev/download and run your system’s installer:

  • macOS: open the .dmg and drag Magnitude to Applications. Choose the Apple Silicon or Intel build.
  • Windows: run the .exe. It installs to %LOCALAPPDATA%\Programs\Magnitude.
  • Linux: install the .deb or .rpm for your distribution, for example with sudo apt install ./magnitude-desktop.deb or sudo dnf install ./magnitude-desktop.rpm.

Then open Discover, pick a recommendation, and press Download. Models and settings are saved in ~/.magnitude (%USERPROFILE%\.magnitude on Windows), separate from the app, so uninstalling won’t delete them.

Some Hacker News users saw the initial “Assessing models” step hang. The team acknowledged it as a bug and is tracking it in issue #142 of the repository.

Can you use Claude Code for free with local models?

Yes, as far as inference goes: connected to Magnitude, Claude Code can use a model running on your machine, with no token charges. Open Connections, search for Claude Code, press Connect, pick a downloaded model, and copy the command the card shows you.

It’s worth knowing what Connect does under the hood:

  • It rewrites your config. It modifies ~/.claude/settings.json so Claude Code routes through Magnitude’s gateway.
  • Local models have their own prefix. Their IDs start with anthropic-local/, so you start Claude Code with claude --model anthropic-local/MODEL_ID. The Connections card already includes the right ID, and on Windows you paste the command in PowerShell.
  • It also affects hosted models. While connected, even requests to cloud models route through the gateway, so Magnitude has to stay open.
  • Undo it before switching back. To use Claude Code normally again, disconnect it in Magnitude and restart Claude Code.

What about OpenCode, Codex, or other agents?

The flow is the same for all: Connections → your agent → Connect, then restart the agent and pick a Magnitude model. Supported agents are Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline. We already covered two of them in Hermes Agent and in our note on OpenClaw (topic 1915).

Who should try Magnitude now?

Try it now if you’re running code agents locally on an M4-class Mac or a single recent NVIDIA GPU, and prefer not to hand-tune llama.cpp parameters. That’s where the vendor figures were measured and where the profiling that “picks the model for you” saves the most time.

Wait, or run your own benchmark first, if any of these apply to you:

  • you’re on an M5-series Mac
  • you work with multiple GPUs or AMD on ROCm
  • you depend on MLX-specific optimizations

The team is transparent about these gaps and the engine changes fast. Base your decision on your own tok/s, not the 2× headline.