// the find
swyxio/gh-action-data-scraping
this shows how to use github actions to do periodic data scraping
A reference implementation, not a library, showing how to turn a GitHub Actions cron job plus a commit-back step into a free periodic scraper with git itself as the datastore. It's aimed at people who want to scrape a site on a schedule without standing up a server or paying for storage, and the repo's own multi-year run of daily JSON snapshots doubles as the proof it works.
The pattern is genuinely minimal: checkout, npm install, run a script, commit the output, done - no framework, no server, nothing to deploy. Storage and scheduler are both free on public repos, which is the actual selling point. The data/ directory isn't a mockup - there are hundreds of real daily snapshot files going back to January 2020, so this isn't a tutorial that was never run past the demo.
The 5-minute minimum cron interval and GitHub's 1GB soft repo limit are hard ceilings from the platform, not something this project works around (LFS is mentioned as a maybe). There's no dedup or pruning logic - every run just adds another file, so a long-lived scraper accumulates unbounded history unless you build that yourself. It's one action.js file and one workflow YAML; there's no abstraction to adopt, just a pattern to copy and adapt by hand. The README itself points out that GitHub's own Flat Data action has covered this use case natively since 2021, so the repo is really a historical writeup at this point.