finds.dev← search

// the find

magnitudedev/magnitude

★ 6,439 · Rust · Apache-2.0 · updated Oct 2026

Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD, or just a CPU.

Magnitude is an open source inference engine and desktop app that compiles and tunes its kernels on the machine it runs on, then serves open-weight models to agents through an OpenAI-compatible API. It is aimed at people running local models on Apple Silicon, NVIDIA, AMD, or CPU-only machines who want more speed without picking kernels or tuning per chip by hand.

- The engine design is written down in design/inference, covering the scheduler, prefix cache, residency, memory, and speculative generation. Most inference projects publish far less about their internals than that.

- Hardware support is tracked as data rather than asserted in prose. assets/hardware has a facts file, an inventory, a coverage.json, and audit scripts, so you can check whether a specific GPU or laptop is covered before you install.

- Integration sits at the API boundary. Anything that speaks the OpenAI-compatible endpoint can use it, so it is not tied to one agent harness.

- The 'up to 2x' headline is a best case. The README's own figures are 92% faster decode on Metal but 19% on CUDA, with prefill gains of 9% and 23%. It does not state the models, quantizations, or context lengths behind the charts, so the numbers are hard to reproduce or compare against your own llama.cpp build.

- The README does not say how long on-device kernel tuning takes on first run, or whether the results are cached across model changes. Since on-device tuning is the core claim, that is the first thing a new user will want to know.

- The visible tree is mostly TypeScript and Electron (cli/ and desktop/, built with Bun and electron-vite), and the engine itself is not in the truncated listing. Anyone who wants to read or embed the Rust engine should expect the JS and Electron layers to be most of what they see first.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →