Case file — 029A0D73
Launch roast · Launch HN
Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents by anerli · HN thread
Unsolicited, from public launch info only. Founder? Reply or ask us to take it down.
The idea
“Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents https://github.com/magnitudedev/magnitude Founder's Show HN post: Hey HN, Anders and Tom here. We're building Magnitude, an inference engine for agents that optimizes itself to run as fast as possible on your hardware. It works on Mac, Linux, and Windows on any hardware and is up to 2x faster than llama.cpp. We're both software engineers and previously built an open source browser agent to 4k+ GH stars and 100k+ downloads. We increasingly wanted to run it on local models, but found that no inference engine worked for our use case. Inference engines today all make a performance tradeoff. They are either: - Built for batched inference on datacenter hardware at the cost of single-session performance (vLLM, SGLang) - Designed for broad compatibility instead of optimizing for specific hardware (llama.cpp, Ollama) - Specialized for specific hardware or models but lacking engine completeness (oMLX, ds4) Plus none of them are designed for running agents locally. Sessions are long, several often run at once, and you still want to use your computer for other things. Magnitude is built for maximum performance on your hardware and running local agents: - On-device compilation and tuning: Kernels are written with flexible parameters that are tuned on your actual device before the model runs. This gives you broad hardware compatibility with the same performance ceiling as hardware-specific kernels. - Focus on best architectures: We write our tunable, highly efficient kernels for the most popular open-weights families. This allows us to achieve and surpass the performance of hardware or model specialized engines, without forcing ourselves to over-generalize at the cost of performance. - Dynamic memory allocation: Magnitude reserves only enough memory up front to hold model weights. As your agent sessions grow, the memory heap dynamically increases, and frees itself when agents stop. Your hardware can still be used for other stuff while agents run. - Hybrid paged attention: We borrow the best ideas from engines like SGLang to allow concu Landing page content (magnitudedev/magnitude): # magnitudedev/magnitude Open source inference server that runs the best local models for your hardware, plugged into the agent you already use. Works with Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline. - Stars: 1704 - Forks: 133 - Watchers: 1704 - Open issues: 13 - License: Apache License 2.0 - Homepage: https://magnitude.dev - Default branch: main - Created: 2026-06-12T09:06:26Z ## Languages - C++ - CMake - CSS - HTML - JavaScript - Jinja - Python - Rust - Svelte - TypeScript ## Top Contributors - anerli (380 contributions) - thrgreenwald (173 contributions) - github-actions[bot] (39 contributions) - fabianhug (1 contributions) - nicolasdmolina (1 contributions) --- ## README Magnitude Run your agent on local models. Free, private, and offline. Magnitude is an open source inference server that runs the best local models for your hardware, plugged into the agent you already use. It profiles your machine, recommends the models that fit, then downloads, tunes, and runs them. Works with Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline, or use the built-in harness. ⭐ Help us reach more developers and grow the Magnitude community. Star this repo! ## Get started **Send this to your agent to walk through models and setup:** ```text Set up local models for me with the Magnitude CLI. Install it with `npm i -g @magnitudedev/cli` (or my package manager), then run `magnitude docs onboarding` and follow the instructions. ``` Your agent will profile your hardware, walk you through the best local models for it, download the ones you pick, and switch itself over to them. Magnitude supports macOS and Linux. Windows is supported through WSL. Want to browse the models directly? ```sh npm i -g @magnitudedev/cli magnitude setup ``` The interactive setup lets you browse the recommended models and choose one yourself. ## Why Magnitude? - **Free to run:** no token costs, API keys, or rate limits - **Fully private and offline:** models, prompts, and files stay on your machine - **Agent-first setup:** one prompt and your agent walks you through the rest - **Knows your hardware:** profiles your chip, memory, and bandwidth - **Recommends what fits:** the best models for your machine, with estimated tok/s - **Tuned end to end:** speculative decoding, concurrency, all set for your machine - **Models on demand:** loaded on request, unloaded when idle or memory fill”