// the find
waybarrios/vllm-mlx
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
An inference server that puts vLLM's continuous batching and paged KV cache on top of MLX, so it runs on Apple Silicon instead of CUDA. It speaks both OpenAI's and Anthropic's HTTP APIs from the same process, which makes it a drop-in local backend for Claude Code, OpenCode, and similar coding CLIs. Aimed at people who want to run LLMs, VLMs, TTS/STT, and embeddings locally on a Mac without juggling four separate servers.
Paged KV cache with prefix sharing and SSD tiering is a real feature, not just a checkbox — spilling prefix cache to disk for long-context agent sessions is something most local-inference tools skip entirely. The reranker code explicitly whitelists supported activation functions and fails loudly on anything else instead of silently running the wrong math, which is the kind of correctness detail that's easy to skip and expensive to debug later. Test suite is large and specific (dedicated files for streaming edge cases, tool-call promotion, KV cache quantization, batching determinism), suggesting the batching and caching internals have actually been exercised rather than assumed to work.
Locked to Apple Silicon via MLX/Metal — there's no path to multi-GPU or multi-node serving, so despite the 'vLLM-style' framing this can't scale past one machine, unlike the project it's modeled on. The cross-CLI compatibility claim in the README is a single run against one model (Qwen3.8-27B-4bit) on one date, not an ongoing compatibility matrix — treat it as 'worked once' rather than 'works reliably across versions.' The feature surface is enormous for one project (LLM serving, vision, audio TTS/STT, embeddings, reranking, MCP tool parsing, speculative decoding, MoE tuning) and the README's citation lists a single author, so it's worth checking the actual contributor graph before betting on long-term maintenance of all of it. No mention of auth or rate limiting on the server itself, so anyone exposing this beyond localhost needs to put something in front of it.