// the find
catalyst-cooperative/pudl
The Public Utility Data Liberation Project provides analysis-ready energy system data to climate advocates, researchers, policymakers, and journalists.
PUDL is a data pipeline that scrapes and normalizes US energy regulatory filings (EIA, FERC, EPA CEMS, PHMSA, and more) into Parquet/DuckDB warehouses with hundreds of cleaned tables. It's for researchers, journalists, and analysts who need utility/emissions data without hand-parsing government spreadsheets themselves.
Raw inputs get archived to Zenodo with DOIs before any processing, so if an agency deletes or reformats a dataset the pipeline still has a frozen source to rerun against — a reproducibility problem most scrapers ignore. The dbt layer has real data quality constraints (ratio checks, weighted quantiles, subcomponent-sum validations) rather than just dumping tables. Nightly automated builds publish to AWS Open Data Registry and Kaggle for free, so you can pull processed output without running the pipeline yourself. CI is mature: pytest on merge groups, codecov, pre-commit.ci, docs build status all tracked.
The dbt/models tree alone has hundreds of subdirectories covering a dozen+ source agencies — onboarding into this codebase to contribute (versus just consuming the output) is a serious investment. The long-standing SQLite output is being deprecated in 2027, so anyone with an existing integration built on pudl.sqlite has a forced migration to Parquet/DuckDB ahead. Coverage is uneven: EIA Form 176 and 191 are explicitly marked work-in-progress while others are fully processed. It's entirely US-focused (EIA/FERC/PHMSA/USDA), so it's a non-starter if you need utility data for any other country.