finds.dev← search

// the find

ijmarshall/cochrane-nlp

★ 17 · Python · GPL-3.0 · updated May 2016

files for systematic review automation project

A research codebase for automating aspects of Cochrane systematic reviews - extracting population sizes, quality assessment, matching Cochrane reviews to PubMed abstracts, and co-training NLP models on that parallel corpus. Aimed at NLP/ML researchers in biomedical text mining, not general-purpose developers.

Includes a genuinely useful parallel corpus linkage (biviewer.py + pickle data) mapping Cochrane review records to their source PubMed abstracts, which is hard to construct yourself. The quality*.py files show iterative experimentation on risk-of-bias classification, and pipeline.py separates feature extraction from the sklearn modeling step cleanly enough to reuse.

No requirements.txt or setup.py, so reproducing the environment (NLTK version, sklearn version, custom Brill tagger pickle) is guesswork. There are five numbered quality*.py files (quality.py through quality5.py) with no explanation of which is current or why the others exist - that's dead research code left in the repo. Last commit is from 2016, no tests, and the README doesn't explain the actual research pipeline end-to-end (how bilearn.py, biviewer.py, and pipeline.py fit together to produce a trained model).

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →