AI News Daily한국어RSS
Friday, October 2, 2026

Story 7 of 7

Magnitude, a local inference engine that compiles kernels for your exact chip, claims up to 2x llama.cpp

First seen on X 20 hours ago@Dontgiveup_26 ♥ 823

Among local AI developers who use Ollama and llama.cpp, an open-source inference engine called Magnitude is making the rounds. Most engines ship kernels precompiled per hardware class, one path for Apple Silicon and one for NVIDIA. Magnitude compiles and tunes its kernels on your device before the first run. It claims up to 2x the speed of llama.cpp, with decode 92% faster on Metal and 19% faster on CUDA, and runs on Apple Silicon, NVIDIA, AMD or a bare CPU.

It is written in Rust and connects the coding agent you already use (Pi, OpenCode, Hermes, Codex and others) with one click. A Korean-language post introducing it drew 800 likes.

More stories from the same day

  1. Claude Code gets mods: a few lines of TypeScript now change how it behaves and looks
  2. Anthropic opens claude.dev, a developer blog with Opus 5.5 playbooks and engineering write-ups
  3. After a day of 'GPT-6.1 Sol is slow' reviews, Altman says it was load and should be much better now
  4. Posts claim Fable 5.1 is quietly serving a 'Fable 5.5', and Anthropic has said nothing
  5. Gemini 4 Argon will open to Google AI Ultra subscribers and paid API customers first
  6. OpenDots, an open-source dots clone you host yourself, is out
See all of Friday, October 2, 2026 →

* Summaries and publishing on this page are done by an automated agent. Check the facts against the original links.

Privacy PolicySupport