finds.dev← search

// the find

AarambhDevHub/aarambh-studio

★ 110 · Rust · Apache-2.0 · updated Aug 2026

🦀 Decoder-only LLM built from scratch in pure Rust using Candle — no Python, no PyTorch. Gated DeltaNet + sparse attention, fine-grained MoE, native video/document understanding, long-horizon tool agents, quantization-aware training. Scales: Tiny (25M) to Large (1.3B).

A from-scratch decoder-only LLM stack in pure Rust on top of Candle — tokenizer, training, quantization, fine-tuning, serving, and multimodal input, with no Python/PyTorch dependency anywhere in the inference path. Aimed at Rust developers who want to understand or hack on an LLM's internals without touching a Python ML stack, not at people who want a model that already works.

The core transformer path (RMSNorm, RoPE, GQA, SwiGLU, KV cache, CPU SIMD attention, optional CUDA kernels) is genuinely implemented rather than stubbed, and doing that without leaning on PyTorch bindings is real work. CI enforces `clippy -D warnings -D clippy::undocumented_unsafe_blocks` and `rustdoc -D missing_docs`, which is a stricter bar than most solo Rust ML projects bother with. The workspace is cleanly split into 20 focused crates (tokenizer, kernel, nn, quant, serve, etc.), and it supports real interop formats (GGUF, SafeTensors, GPTQ/AWQ) instead of a made-up checkpoint format nobody else can load.

The README lists 55 numbered 'phases' covering MLA, native audio, sparse MoE, multi-node training, RLAIF, sandboxed multi-agent orchestration, from-scratch RAG, model merging, and auto-generated red-team model cards — that's a research-lab roadmap compressed into a 110-star repo, and breadth at this scale is a classic sign the implementations are thin layers that compile and pass a smoke test rather than code that's been stressed under real training runs. No pretrained weights ship and no crates.io publish exists, so there's no way to check whether any of this actually produces a usable model — you're trusting the code on faith. The capable paths are gated behind CUDA (sparse MoE dispatch, custom PTX kernels) while the CPU path — the one actually differentiating this from 'just use PyTorch' — gets the degraded dense fallback. The phase-numbered doc structure and version-locked architecture docs per release read like a checklist driven by an LLM coding assistant rather than organic iteration against real bottlenecks, which is worth confirming before trusting any of the more exotic components (RLAIF, test-time scaling, model merging).

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →