// the find
NVIDIA/Model-Optimizer
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.
A single library from NVIDIA that bundles post-training quantization, QAT/QAD, pruning, distillation, speculative decoding, and sparsity into one API, with export paths into TensorRT-LLM, vLLM, and SGLang. It's for teams deploying large HF/PyTorch/ONNX models (mostly LLMs/VLMs) who need to shrink them for inference rather than researchers experimenting with novel compression theory.
Covers the whole compression pipeline under one package instead of stitching together separate quantization, pruning, and distillation libraries with incompatible checkpoint formats. The export path is real, not academic — a unified HF export API feeds directly into TensorRT-LLM/vLLM/SGLang, and the published Nemotron/DeepSeek-R1/Llama checkpoints on Hugging Face prove the pipeline actually produces deployable artifacts with reported throughput and accuracy numbers. NVFP4 support for Blackwell is ahead of most competing quantization toolkits.
Value is concentrated on NVIDIA hardware and NVIDIA's own inference stack — outside that ecosystem (AMD, CPU-only, non-NVIDIA accelerators) a lot of the headline features (NVFP4, Blackwell speedups) just don't apply. Still pre-1.0 with an explicit policy that breaking changes can land in minor version bumps, so pin versions carefully. The example tree is enormous and fragmented across frameworks (Megatron-Bridge, llama_factory, diffusers, ONNX, Windows, Alpamayo), which makes it hard to find the one path relevant to your model without reading several READMEs first. The top-level README itself is mostly a rolling blog/announcement feed rather than a technical overview, so you're pushed into the docs site before you can tell if a given technique fits your use case.