finds.dev← search

// the find

JustVugg/colibri

★ 37,962 · C · Apache-2.0 · updated Sep 2026

Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

A pure-C inference engine that treats VRAM, RAM, and NVMe as one memory hierarchy so multi-hundred-billion to multi-trillion parameter MoE models can run by streaming inactive experts off disk instead of holding the whole model resident. It's for people who'd rather buy a fast SSD and build from source than rent GPU time, and who are fine with single-digit-tokens-per-second throughput in exchange for holding a huge model on hardware they own.

The expert-placement mechanism is a specific, well-described technique (measured routing heat driving a per-layer LRU plus one-layer-ahead prefetch reported at 71.6% accuracy) rather than just 'quantize and mmap.' The benchmark docs report negative and neutral results next to the wins — O_DIRECT called out as drive-dependent, a third slower SSD found neutral after striping, MTP speculative decoding shown losing 32% at 85% expert hit — which is more honest than most projects bother to be. Persisted KV state and prefix checkpointing for warm restarts is a genuinely useful detail for anyone running repeated or long agent sessions against a model that's slow to cold-start. Nine model families sharing one CLI front end (chat/serve/web) with per-family C engines is a clean way to keep the core small while still covering real architectures.

Every headline throughput number is low single-digit tokens/second (0.05-0.1 tok/s cold on a 25GB box, 5.8-6.8 tok/s even on 6x RTX 5090s) — the README's own numbers undercut the 'run frontier models on hardware you own' framing; this is a research platform for poking at huge models locally, not something you'd serve traffic from. Getting started means either a 372GB-1.6TB per-model download or a manual multi-hour FP8-to-int4 conversion, plus a separate `make -C c <family>` build per model — there's no single binary that works across models out of the box. The correctness story rests on a transformers oracle matching 'typically 30-32/32' tokens with two unexplained floating-point near-ties left open, which is the right instinct but leaves an unresolved edge case that could bite on an unlucky prompt. The README itself leans heavily on dashboards, a 'Brain' cortex visualization, and named modes ('Brio') that read like the same over-selling the project claims to avoid elsewhere — the real signal is in docs/benchmarks.md, not the landing-page framing up top.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →