finds.dev← search

// the find

ggml-org/whisper.cpp

★ 54,158 · C++ · MIT · updated Oct 2026

Port of OpenAI's Whisper model in C/C++

whisper.cpp is a C/C++ port of OpenAI's Whisper speech recognition model, built on the ggml tensor library with no external dependencies. It is aimed at developers who need offline transcription inside an app, on a server, or on hardware where shipping a Python stack is not an option.

- The model's high-level logic lives in whisper.h and whisper.cpp, so the whole forward pass can be read in one sitting. For a model this size that is rare, and it makes the code a useful reference for anyone porting transformer inference.

- Hardware coverage is wide for one codebase: Metal, Core ML and ANEForge on Apple Silicon, CUDA, Vulkan, ROCm, OpenVINO, CANN, MUSA, and a VitisAI path for AMD NPUs, all behind CMake flags with a CPU fallback.

- Integer quantization through the quantize tool shrinks the checkpoints considerably. The ggml format packs weights, mel filters and vocabulary into one file, so there is no separate tokenizer or Python runtime to deploy.

- The C API plus the language bindings (Go, Rust, Ruby, Java, .NET, JavaScript and WASM, Python) make it embeddable outside C++. The server, stream and command examples cover the common integration shapes.

- The README says it plainly: inference only. There is no training or fine-tuning here, so adapting the model to a domain vocabulary or accent means going somewhere else.

- The accelerator matrix is large and not every path gets equal testing. Core ML, OpenVINO and VitisAI each need device-specific setup, the first run on a device compiles slowly, and the cache files are tied to the machine and OS build.

- The CLI reads only 16-bit WAV unless it is built with FFmpeg, so most real input needs a conversion step first. The large checkpoints need about 3.9 GB of RAM and 2.9 GiB of disk, which rules out small devices.

- Word-level timestamps via -ml 1 and the tinydiarize speaker turns are labelled experimental, and the microphone example is described as naive. Measure them on your own audio before depending on them.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →