// the find
skywind3000/ECDICT
Free English to Chinese Dictionary Database
ECDICT is a free, open English-to-Chinese dictionary dataset (CSV/SQLite/MySQL) with roughly 770k+ entries, plus a small Python toolkit (stardict.py) for reading, converting, and querying it. It's aimed at developers building dictionary apps, Anki deck generators, vocabulary-learning tools, or NLP preprocessing for Chinese learners of English, not at building a dictionary from scratch.
The exchange/lemma data (tense, plural, comparative forms, and lemma mapping) derived from BNC plus WordNet/NodeBox is something most free dictionary datasets don't bother to include, and it's genuinely useful for stemming raw text. Dual frequency ranking (historical BNC vs. a contemporary corpus) is a smart touch — it lets you tell apart words that are archaic-but-common historically (like 'quay') from ones trending now. Storing everything as CSV keeps PRs and diffs reviewable on GitHub, and stardict.py gives one consistent query interface whether you're backed by CSV, SQLite, or MySQL. Exam-vocabulary tagging (CET4/6, IELTS, Oxford 3000, Collins star rating) is a nice freebie if you're building a leveled vocabulary tool.
There's no test suite, no CI, and it's not packaged for PyPI — stardict.py is a script you vendor into your project, not a library you pip install. Documentation is Chinese-only with no English README, which will block a chunk of potential adopters outright. Data quality is crowd-sourced and self-reported, with accuracy fixes tracked only informally in a HISTORY section rather than real changelogs or data provenance per entry. The 'detail' (example sentences) and 'audio' fields have been marked as TODO/not-yet-added since the schema was first documented, and the full dataset ships separately as a 7z archive outside the normal CSV diff workflow, which undercuts the 'easy to review and PR' story for anything beyond the base 76-万-entry CSV.