// the find
DataExpert-io/data-engineer-handbook
This is a repo with links to everything you'd ever want to learn about data engineering
A link-aggregation repo for breaking into data engineering, bundled with some real bootcamp materials (SQL, PySpark, Flink labs with homework and tests). Good for total beginners who want a reading list and a few hands-on labs; less useful for anyone past the fundamentals stage.
The bootcamp materials folders are actual working code, not just slides — docker-compose setups, PySpark jobs with pytest tests, a Flink streaming job with init SQL. The design patterns links (cumulative table design, microbatch dedup) point to genuinely solid technical write-ups, not generic blog fluff. High star count and an active push history show the content gets kept current.
Most of the README is promotional: follower counts for the author's own YouTube/LinkedIn/TikTok accounts, discount codes for paid bootcamps, cross-links to the author's company products. The actual handbook content is a flat list of external links with zero evaluation — no indication of which orchestrator or warehouse is actually good versus which one is there because they sponsor something. Jupyter Notebook as the primary language tag is misleading; the repo is mostly markdown and SQL, notebooks are a minor fraction.