⬢github Rust · 6.1K ★ +609 since we first saw it · pushed 8 min ago · Apache-2.0
magnitudedev/magnitude
Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD, or just a CPU.
Magnitude is an open source (Apache 2.0) LLM inference engine, shipped as a desktop app plus CLI, that compiles and tunes its GPU kernels on your specific hardware before running a model. It claims up to 2x faster decoding than llama.cpp, lower memory use, shared prefix caches for concurrent agent sessions, and one-click integration with agents like Pi, OpenCode, Hermes, and Codex.
Why now: It just launched on Hacker News as a YC S25 launch, staking a bold performance claim of beating llama.cpp on both Metal and CUDA, which is drawing attention and comparisons.
Who it is for: Developers running open-weight models locally to power coding agents on Apple Silicon, NVIDIA, AMD, or CPU-only machines.
Stars over our 111 snapshots: 5.5K to 6.1K, since 1 d ago.
Where people talked about it
API: https://socialmediatrends-api.osmike.com/v1/repos/magnitudedev/magnitude