finds.dev← search

// the find

QuentinFuxa/WhisperLiveKit

★ 11,096 · Python · Apache-2.0 · updated Sep 2026

Real-time, local speech-to-text with streaming ASR, speaker diarization, translation, and OpenAI/Deepgram-compatible APIs.

A self-hosted real-time speech-to-text server built around actual streaming ASR research (AlignAtt/SimulStreaming, LocalAgreement) instead of just re-running Whisper on rolling audio chunks. It bundles diarization, translation, and OpenAI/Deepgram-compatible endpoints, aimed at people who want live transcription without sending audio to a cloud API.

The core problem it solves is real: naive chunked Whisper mangles words at chunk boundaries, and this actually implements published incremental-decoding policies to fix that, with a causal-KV Qwen3-ASR backend that gets constant per-chunk compute instead of re-encoding a growing window. It ships a real benchmark suite (scatter plots of WER vs RTF on H100, reproducible via a script) rather than just claiming speed. Backend breadth is unusually wide — Whisper variants, Voxtral, Qwen3-ASR, Canary, FunASR — behind one CLI and one WebSocket protocol, and the OpenAI/Deepgram-compatible surface means it can be a drop-in replacement in existing client code.

The dependency matrix is a minefield: extras like qwen3-vllm, voxtral-hf, canary, and qwen3-vllm-metal explicitly conflict with each other and need separate environments, so picking a backend combo is trial-and-error against pyproject.toml. Diarization is architecturally limited to one speaker per frame even though Sortformer emits overlapping-speech probabilities, so overlapping talkers just get assigned to whichever speaker wins. Several backends are self-described as experimental (vLLM Metal/CUDA causal mode) or partially broken (--disable-punctuation-split is documented as non-functional), and most of the accuracy/latency numbers are H100-only — CPU and non-Apple-Silicon Mac users are mostly flying blind on real-world performance.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →