// the find
SkalskiP/vlms-zero-to-hero
This series will take you on a journey from the fundamentals of NLP and Computer Vision to the cutting edge of Vision-Language Models.
A planned curriculum walking from Word2Vec and CNNs up through CLIP and modern VLMs like LLaVA and PaliGemma, built around paper-by-paper notebooks. Aimed at people who already know some ML and want implementation-level understanding of the architectures that led to vision-language models, not just API usage.
The paper selection is genuinely good and well-sequenced — it traces the actual lineage (word2vec to seq2seq to attention to ViT to CLIP to LLaVA) rather than just listing buzzwords. From Pietro Skalski, who has a track record of readable, implementation-first ML content (roboflow/supervision), so the pedigree is real even if this particular repo isn't there yet.
This is vaporware right now: the README literally says 'coming: january 2025', 5 of 6 topic folders contain nothing but a .gitkeep, and exactly one notebook exists out of a roadmap listing ~20 papers. Last push was January 2025 with no commits since, so whatever momentum existed appears to have stalled. The 1182 stars are pure anticipation/name-recognition, not a signal the content delivers. Nothing here to actually run or learn from beyond the single Word2Vec notebook — don't feature this as if it's a finished resource.