// the find
abetlen/llama-cpp-python
Python bindings for llama.cpp
Python bindings for llama.cpp with three layers: a ctypes wrapper over the C API, a high-level Llama class for completions and chat, and an OpenAI-compatible web server. It is for Python developers who want to run GGUF models locally from application code, or who want a local endpoint that existing OpenAI client code can point at.
The low-level layer is a direct ctypes mirror of llama.h, so new upstream functions are reachable without waiting for a wrapper release. The high-level Llama class does more than template substitution: it has named chat formats, a response_format JSON schema mode that constrains output, speculative decoding through a prompt-lookup draft model, and handlers for LLaVA-style multimodal models. Install is more flexible than most bindings: CMAKE_ARGS passes any llama.cpp backend flag through, and pre-built CUDA, Metal, ROCm and Vulkan wheels exist for people who would rather not compile.
The default pip install compiles llama.cpp from source, so a machine without a working C/C++ toolchain fails at install time, and the Windows and Apple Silicon sections of the README exist because people hit this. Every upstream llama.h change needs hand edits to llama_cpp/llama_cpp.py, and new model architectures only arrive when the vendored llama.cpp submodule is bumped, so the package trails upstream and upgrades can break. Pre-built wheels are pinned to specific CUDA versions, GPU compute capabilities and Python versions, and that matrix is spread across the README, so confirming your hardware is covered takes effort. The project began as one author's personal tool by its own account, so pin your version and expect to re-test on each upgrade.