// the find
hadley/stats337
Readings in applied data science
This is the reading list and syllabus for a one-off Stanford discussion seminar Hadley Wickham taught in Spring 2018, not a software project — despite the R language tag there's no code here, just a README of links to papers and blog posts on the non-technical side of data science (collaboration, reproducibility, ethics, careers). It's for someone who wants a reading list on the 'soft' parts of doing data science, not for anyone looking for a library or tool.
The link selection is genuinely good — mostly free-to-read blog posts and open-access papers rather than paywalled journal articles, so you can actually click through and read them. Topic breadth is wide and deliberately underserved: version control for non-engineers, spreadsheet hygiene, DevOps for analysis, and a full week on ethics, areas most 'intro to data science' lists skip entirely. The student annotated bibliographies are a nice bonus — they extend the list well past what's in the README itself.
Dead since June 2018 with no code, no CI, nothing to run — the repo is just a markdown file, so 'stars: 1611' is really bookmarking behavior, not adoption. A chunk of the linked content has surely rotted or moved (Google Docs, old blog domains, a Twitter thread as a citation), and there's no archive or wayback-link fallback. Several of the annotated bibliographies are PDFs, so they're not searchable or linkable at the section level. It's also explicitly built around one Stanford course's grading rubric (percentages, due dates, check/plus-minus grading), which is dead weight if you're just here for the reading list.