Developers running agents locally have long faced a structural gap. Engines like vLLM and SGLang are built for datacenter batch inference, sacrificing single-session performance; llama.cpp and Ollama pursue broad compatibility instead of squeezing the maximum out of any specific chip; specialized engines like oMLX and ds4 lack ecosystem completeness. In other words, almost no engine has been seriously designed for the scenario of running several long agent sessions on your own machine at once. On September 30, a YC S25 team released Magnitude, an open-source inference engine, on Hacker News. Its idea: move performance tuning onto your own device.

What it does differently

Magnitude is written in Rust, ships its own GPU kernel runtime and autotuner, and is open source under Apache 2.0. The team lays out three technical moves. First, kernels carry flexible parameters and are compiled and tuned on your actual device before a model runs, letting generic kernels reach the performance ceiling of hardware-specific ones. Second, memory allocation is dynamic: at startup only enough memory for model weights is reserved; the heap grows as agent sessions grow and frees itself when agents stop, so the machine stays usable for other work. Third, hybrid paged attention borrows the radix-cache idea from SGLang so concurrent sessions share prefix caches, while placement is optimized for memory adjacency so single-session performance does not collapse under concurrency. The README lists FlashAttention, FlashInfer, and TurboQuant among its inspirations.

The vendor-reported numbers

Against llama.cpp, with a unified configuration of Qwen 3.6 35B A3B (4-bit), 64k context, and speculative decoding disabled: on a Mac M4 Pro 48GB, decode is 92% faster (30 to 57 tok/s), prefill 9% faster (466 to 507 tok/s), and per-agent memory 28% lower; on a DGX Spark, decode is 19% faster (49 to 58 tok/s), prefill 23% faster (2033 to 2507 tok/s), and memory 27% lower. The product ships as a desktop app that connects Pi, OpenCode, Hermes, Codex, Claude Code and other clients with one click, with an OpenAI-compatible API for everything else; models spin up on demand when agents need them and shut down after inactivity. About a day after launch, the HN post sat at 187 points with 88 comments, and the repository at roughly 6k stars.

Discount the self-reported numbers accordingly

The angle deserves credit: single-session latency and multiple co-resident agents are real pain points of local inference — datacenter engines do not bother optimizing for them, and generalist engines are not targeting that objective. But every benchmark here is vendor-tested: a single model, a single quantization, speculative decoding off — exactly the configuration interval that favors the vendor. The up-to-2x tagline takes the best Metal decode number, while the CUDA side is 19%. Until third-party replication appears, the safer reading is a well-intentioned directional validation, not a settled verdict. The roadmap's expert streaming — parking MoE experts in RAM or on disk and loading them into the GPU just in time — is worth watching; that is the real key to running large models on small VRAM.

For developers who keep agents running locally, this adds a free and open-source option; for the inference-engine landscape, the local-engine-for-the-agent-era niche is being filled fast. A question to leave with you: when agents truly live on your machine, is your current inference engine actually optimized for that?

References: official launch post https://news.ycombinator.com/item?id=49911995 ; GitHub repository https://github.com/magnitudedev/magnitude