finds.dev← search

// the find

turboderp-org/exllamav3

★ 1,562 · Python · MIT · updated Sep 2026

An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs

ExLlamaV3 is a quantization and inference library for running LLMs on one or two consumer GPUs, built around its own EXL3 format (a QTIP/QuIP#-derived method) plus tensor/expert-parallel inference and CPU offloading for MoE models. It's for people already running local LLMs who want better quality-per-bit than GPTQ/AWQ and are comfortable dealing with a from-source CUDA build.

EXL3 isn't another round-to-nearest quantizer — it computes Hessians on the fly and uses a fused Viterbi kernel, so a single conversion run (minutes for small models, hours for 70B+ on one card) gets real quality gains at low bitrates, not just smaller files. Architecture coverage is large and current: dozens of HF model classes including fresh releases with vision/multimodal variants, not just the usual Llama/Mistral set. CPU offloading with AVX2/AVX512 paths lets big MoE models run split across limited VRAM and system RAM instead of requiring everything to fit on the card. The custom CUDA kernel set — per-bitrate GEMV/GEMM, hadamard transforms, fused attention, parallel all-reduce for both CPU and GPU — is genuine low-level engineering, not a wrapper around someone else's kernels.

Getting it running is real work: you need a torch build matched to one of six CUDA flavors, a prebuilt wheel tied to a specific torch/python/CUDA combo, or a from-source build with a JIT compile that takes several minutes on first import — the PyPI package ships no prebuilt extension at all. NVIDIA/CUDA only, nothing for AMD/ROCm, despite the "consumer-class GPU" framing. Windows support feels bolted on rather than first-class (needs a separate triton-windows package). It's effectively a single-maintainer project carrying dozens of architecture-specific files that each need updating as upstream model code shifts, which is a lot of surface area for one person to keep in sync with every new model drop.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →