finds.dev← search

// the find

BlankerL/DXY-COVID-19-Data

★ 2,173 · Python · MIT · updated Jun 2025

2019新型冠状病毒疫情时间序列数据仓库 | COVID-19/2019-nCoV Infection Time Series Data Warehouse

A time-series archive of COVID-19 case data for China and worldwide, scraped daily from DXY (Chinese medical portal) starting in early 2020. It's for researchers and analysts who need the raw historical case/death/recovery numbers rather than a live dashboard, but it's now archived and frozen.

Provides plain CSV/JSON files instead of forcing consumers to hit an API, which is exactly what non-programmer researchers asked for. Data was collected consistently on a daily cron from source going back to the outbreak's start, giving a long uninterrupted historical run. The maintainer documented known data quality issues directly in the README (duplicate city counts in Henan, swapped Jilin city numbers) rather than silently leaving them for users to discover.

Project is explicitly archived as of the source going dark in mid-2020s era updates stopped — no new data will ever be added, so it's only useful for historical analysis, not current tracking. The known data errors (duplicated counts, swapped provincial numbers) are documented but never corrected in the actual files, meaning anyone using this for research has to hand-clean it first. Source data was manually entered by DXY staff with no validation layer, so there's no guarantee of accuracy beyond what's already been flagged in issues. Single-maintainer project with no contribution process for fixing the data itself — the README explicitly says customization requests won't be accepted.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →