// the find
ninehills/chatglm-openai-api
Provide OpenAI style API for ChatGLM-6B and Chinese Embeddings Model
A thin FastAPI-style wrapper that exposes ChatGLM-6B/ChatGLM2-6B and a Chinese embeddings model behind an OpenAI-compatible /v1/chat/completions endpoint. Useful for anyone who wants to point existing OpenAI SDK clients (chatbot-ui, OpenCat) at a self-hosted Chinese LLM without rewriting client code.
The OpenAI-shaped API is the actual value here — swap the base URL and API key and existing OpenAI client code works unmodified. It supports int4/int8 quantized checkpoints so it runs on modest GPUs, and it bakes in both ngrok and Cloudflare tunnel support for exposing a home GPU box publicly, which saves the usual reverse-proxy fiddling. Multi-GPU inference via CUDA_VISIBLE_DEVICES is also wired up.
Last commit is from July 2023 and it's tied to ChatGLM-6B/ChatGLM2-6B, both long superseded by GLM-4 and friends — anyone picking this up now is deploying an obsolete model stack. The README is entirely in Chinese with no English documentation, which will stop most non-Chinese-reading developers before they start. There's no test suite, no batching or concurrency handling visible in the API layer, and no auth beyond a single static bearer token in a TOML file — fine for a personal tunnel, not for anything shared.