// the find
explosion/spaCy
💫 Industrial-strength Natural Language Processing (NLP) in Python
spaCy is a production-oriented NLP library for Python covering tokenization, tagging, parsing, NER, and text classification across 70+ languages, with first-class support for plugging in pretrained transformers. It's aimed at people building NLP into an actual application or pipeline, not researchers prototyping new architectures.
The Cython core makes the tokenizer and pipeline genuinely fast compared to pure-Python NLP libraries, which matters once you're processing real volumes of text. The config-driven training system (spacy train with a .cfg file) makes runs reproducible and versionable instead of scattered CLI flags or notebook state. Component architecture (nlp.add_pipe) is clean and lets you swap or disable pipeline stages without touching the rest of the pipeline. Model packaging turns trained pipelines into installable Python packages, which removes a whole class of "where do I load this model from" deployment pain.
Trained pipelines are large binary downloads pinned to specific spaCy versions, and mismatches between installed spaCy and model version are a recurring source of confusing errors. It's a heavy dependency for simple use cases — if you just need basic tokenization or regex-adjacent text processing, this is a lot of machinery. The shift to transformer-backed pipelines in v3 means GPU/CUDA setup complexity leaks into what used to be a CPU-friendly library, and the docs don't always make clear which pipeline size/tier you actually need. Custom component development requires understanding Doc/Span/Token internals fairly deeply before you can extend the pipeline confidently.