finds.dev← search

// the find

neuml/paperetl

★ 699 · Python · Apache-2.0 · updated Dec 2025

📄 ⚙️ ETL processes for medical and scientific papers

paperetl turns messy medical/scientific paper formats (PDF via GROBID, PubMed XML, ArXiv XML, TEI, CSV) into a normalized SQLite/JSON/YAML/Elasticsearch dataset. It's built by neuml, the same team behind txtai, and reads like infrastructure for their own CORD-19-style corpus work rather than a general-purpose tool. Good fit if you're building a search or RAG pipeline over a pile of scientific literature and don't want to write four separate parsers.

Multi-format ingestion (PDF/PubMed/ArXiv/TEI/CSV) is handled through one factory dispatch instead of bespoke scripts per source. Output targets are pluggable via simple URL-style strings (json://, yaml://, or a plain path for SQLite), so swapping destinations is a one-line change. Actually has a real test suite (file database, export, elastic, process) with CI and a Coveralls badge, not just a README full of promises. Codebase is small enough to read end to end in an afternoon if you need to add a parser for a format they don't cover.

PDF parsing isn't self-contained — it depends on a separately-run GROBID Java service, and the README outright admits you'll hit 503s from pool exhaustion that you have to tune around manually. The article schema (schema/article.py) looks hard-coded to their own metadata shape; if your papers don't match it you're patching their code, not configuring an option. Nothing in the docs addresses incremental updates — looks like a directory-to-database wholesale run each time, with no mention of skipping already-processed files. Documentation is one README and a single Colab notebook; there's no reference for writing a new source parser beyond reading the existing ones.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →