Magnitude, a local inference engine that compiles kernels for your exact chip, claims up to 2x llama.cpp
First seen on X 20 hours ago@Dontgiveup_26 ♥ 823
Among local AI developers who use Ollama and llama.cpp, an open-source inference engine called Magnitude is making the rounds. Most engines ship kernels precompiled per hardware class, one path for Apple Silicon and one for NVIDIA. Magnitude compiles and tunes its kernels on your device before the first run. It claims up to 2x the speed of llama.cpp, with decode 92% faster on Metal and 19% faster on CUDA, and runs on Apple Silicon, NVIDIA, AMD or a bare CPU.
It is written in Rust and connects the coding agent you already use (Pi, OpenCode, Hermes, Codex and others) with one click. A Korean-language post introducing it drew 800 likes.