// the find
yihong0618/xiaogpt
Play ChatGPT and other LLM with Xiaomi AI Speaker
A CLI/Docker tool that hijacks a Xiaomi AI speaker's wake word to route your voice queries to ChatGPT or one of a dozen other LLM backends (Gemini, Qwen, Moonshot, Llama3 via Groq, ChatGLM, Doubao), then speaks the answer back through the device or a swappable TTS engine. Aimed at hobbyists who own Xiaomi hardware and want a cheap smart-speaker-to-LLM bridge without rooting the device.
Bot and TTS backends are both pulled out into separate modules (xiaogpt/bot/*.py, xiaogpt/tts/*.py) behind a common base class, so adding a new provider is a contained change, not a rewrite - and the project has clearly been doing this repeatedly (nine+ bot backends, seven TTS options via the tetos library). Config precedence is sane and explicit (cli args > default > yaml/json config), and streaming responses are supported for the APIs that offer them, which matters a lot for a voice interface where latency reads as broken. The Docker setup persists the Mi auth token to a mounted volume with a documented re-auth path for 2FA, which is the part of this stack most likely to require manual intervention.
The whole thing sits on an unofficial, reverse-engineered Xiaomi API (via the separate MiService project) - the README itself documents login failures from risk control for overseas accounts, cookie expiry, and a workaround of copying .mi.token between machines. That's not a dependency risk, that's the core integration being adversarial to its own platform, and it will break again without notice. Hardware support is uneven: certain speaker models (LX04, X10A, L05B/L05C) need a different control mode (--use_command) and lose access to third-party TTS, so the actual compatibility matrix is smaller than the device list implies. Test coverage is essentially nonexistent for the part that matters - tests/ only checks dependency compatibility and the miservice wrapper, nothing exercises the bot or TTS logic. Documentation is Chinese-only with no English README, which will filter out most people landing here from an English search despite the topic having broad appeal.