// the find
mrdbourke/your-first-kaggle-submission
How to perform an exploratory data analysis on the Kaggle Titanic dataset and make a submission to the leaderboard.
A single Jupyter notebook that walks through exploratory data analysis on the Kaggle Titanic dataset and produces a leaderboard submission. It's aimed squarely at people who have never made a Kaggle submission before and want a guided first pass at the classic beginner competition.
The notebook is self-contained — train/test/gender_submission CSVs ship in the repo, so there's no setup friction before you can run it end to end. It pairs the code with a YouTube walkthrough and a Towards Data Science blog post, so you get the reasoning behind each EDA step, not just cells to execute. Scope is tight: one dataset, one goal (get a submission.csv out the door), which makes it easy to finish in one sitting.
There's no requirements.txt or pinned environment, so reproducing the exact pandas/sklearn/catboost behavior years later means guessing versions. It's one notebook, not reusable code — nothing here generalizes past Titanic. No CI, no tests, and the author asks bug reports to be emailed directly instead of filed as issues, which is a dead end for most people. Last pushed mid-2024, so it's a static artifact rather than something that gets maintained.