// the find
supertone-oss-archive/supertonic
Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
Supertonic is a 99M-parameter multilingual TTS model that runs on-device via ONNX Runtime, with example code across 11 languages/platforms (Python, Node, browser/WebGPU, Java, C++, C#, Go, Swift, iOS, Rust, Flutter). It's aimed at people who want offline, no-GPU speech synthesis embedded in an app rather than an API call to ElevenLabs or OpenAI. As of now it's archived — Supertone has ended development, support, and issue triage, so what you see is what you get.
The model is genuinely small (99M params vs 0.7B-2B for comparable open TTS) which means fast cold starts and a real shot at running on a Raspberry Pi or e-reader, not just a beefy laptop. WER/CER numbers on the Minimax-MLS-test benchmark hold up against much larger models like VoxCPM2 across most of the 31 supported languages. Text normalization for things like phone numbers, currency, and unit abbreviations is handled correctly out of the box, which is a common failure point for smaller TTS systems. Having working example code in 11 different runtimes is unusually thorough and saves real integration time.
It's archived: no bug fixes, no security patches, no PR review, so any ONNX Runtime CVE or platform breakage (new Xcode, new .NET, new Node) is permanently on you to fix. There's no built-in voice cloning — that's gated behind Voice Builder, a separate hosted service outside this repo, so 'on-device' really means 'on-device with a fixed preset voice.' The natural-text comparison table (financial expressions, phone numbers) is self-reported with Google Drive audio links rather than a reproducible eval harness, so treat those wins as marketing until you test your own inputs. A few languages regressed noticeably going from Supertonic 2 to 3 (Finnish jumped from 2.29 to 5.40 WER, Vietnamese from 1.48 to 4.49), which the README doesn't call out despite showing the numbers.