// the find
MorvanZhou/Reinforcement-learning-with-tensorflow
Simple Reinforcement learning tutorials, 莫烦Python 中文AI教学
A repo of standalone, single-file TensorFlow implementations of classic RL algorithms — Q-learning through PPO, DDPG, A3C — built as a teaching companion to Morvan Zhou's video series. It's for someone who wants to see the bare mechanics of an algorithm in ~150 lines rather than plug into a maintained RL library.
Each algorithm is isolated in its own file with no shared framework to trace through, so you can read DQN or PPO start to finish in one sitting. The progression from tabular Q-learning/Sarsa up through policy gradients and actor-critic methods is well sequenced for building intuition. Code maps closely to the math, which is the point of a teaching repo.
Dead since March 2024 and built on TF1-era APIs (session-based graphs, placeholders) that don't match how TensorFlow is written now, so getting it running means fighting compatibility issues before you learn anything. No tests, no requirements.txt or dependency pinning, and no docstrings beyond what's in the linked videos. This is a snapshot of 2017-2018 RL teaching code, not something to build on — use Stable-Baselines3 or CleanRL if you actually want to run these algorithms.