finds.dev← search

// the find

zackees/transcribe-anything

★ 1,405 · Python · MIT · updated Sep 2026

Multi-backend whisper app. Blazing fast. Mac-arm optimized. Easy install. Input a local file or url and this service will transcribe it using Whisper AI. Completely private and Free 🤯🤯🤯

A CLI and Python wrapper that routes audio/video to whichever Whisper backend fits your hardware — openai-whisper, insanely-fast-whisper, WhisperX, SenseVoice, MLX for Apple Silicon, or Intel XPU — each built into its own isolated venv on first run. It's for anyone who wants transcription + speaker diarization without hand-rolling CUDA/torch environments themselves, from a one-off YouTube transcript to a GPU box running batch jobs via the new daemon mode.

The isolated-venv-per-backend design is the actual value proposition and it's done properly: each backend gets pinned deps so a WhisperX install doesn't fight an insanely-fast-whisper install's torch version. The daemon/serve mode is a real engineering response to a real problem — cold-start costs (model load, CUDA init) dominate per-invocation time, so paying them once and serving over HTTP is the right call, and it correctly locks backend/token config at startup so clients can't quietly change what they're billing for. CI is unusually thorough for a project this size — separate workflows per OS plus a dedicated WhisperX test job — which matters given how many native dependency combinations this project juggles.

The isolated-env approach has a real cost the README undersells: first run of each backend can pull ~10GB, and the 4.0 cache-path migration orphaned every existing user's venv cache, forcing a silent re-download on upgrade. There was a credential-handling bug where `--hf_token` leaked into stderr and exception tracebacks on subprocess failure — fixed now, but anyone who ran it on a serverless host before the fix needs to rotate tokens, which says something about how long that path went unaudited. Six backends across five OS/hardware combinations is a lot of surface area for what looks like a mostly solo-maintained project; several backends (MLX, XPU, Nix packaging) came from community PRs rather than the maintainer, which is good for breadth but a risk for long-term consistency and support. The README itself is a warning sign: a sponsor ad embedded mid-document, heavy emoji, and marketing phrasing ('blazing fast', 'state of the art') make it harder to tell signal from noise when you're trying to figure out what a flag actually does.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →