// the find
turboderp-org/exllamav3
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs
ExLlamaV3 is a quantization and inference library for running LLMs on one or two consumer GPUs, built around its own EXL3 format (a QTIP/QuIP#-derived method) plus tensor/expert-parallel inference and CPU offloading for MoE models. It's for people already running local LLMs who want better quality-per-bit than GPTQ/AWQ and are comfortable dealing with a from-source CUDA build.
EXL3 isn't another round-to-nearest quantizer — it computes Hessians on the fly and uses a fused Viterbi kernel, so a single conversion run (minutes for small models, hours for 70B+ on one card) gets real quality gains at low bitrates, not just smaller files. Architecture coverage is large and current: dozens of HF model classes including fresh releases with vision/multimodal variants, not just the usual Llama/Mistral set. CPU offloading with AVX2/AVX512 paths lets big MoE models run split across limited VRAM and system RAM instead of requiring everything to fit on the card. The custom CUDA kernel set — per-bitrate GEMV/GEMM, hadamard transforms, fused attention, parallel all-reduce for both CPU and GPU — is genuine low-level engineering, not a wrapper around someone else's kernels.
Getting it running is real work: you need a torch build matched to one of six CUDA flavors, a prebuilt wheel tied to a specific torch/python/CUDA combo, or a from-source build with a JIT compile that takes several minutes on first import — the PyPI package ships no prebuilt extension at all. NVIDIA/CUDA only, nothing for AMD/ROCm, despite the "consumer-class GPU" framing. Windows support feels bolted on rather than first-class (needs a separate triton-windows package). It's effectively a single-maintainer project carrying dozens of architecture-specific files that each need updating as upstream model code shifts, which is a lot of surface area for one person to keep in sync with every new model drop.