// the find
jcjohnson/pytorch-examples
Simple examples to introduce PyTorch
A sequence of short scripts that fit the same two-layer ReLU network to random data, starting from hand-written numpy backprop and climbing through raw tensors, autograd, nn.Sequential, optim, custom Modules, and dynamic graphs. It is for a Python developer who knows the calculus and wants to see what each PyTorch abstraction takes off their hands. The prose targets PyTorch 0.4 and the last push was February 2022, so read it as a conceptual walkthrough rather than current practice.
The ladder is the point: every script trains an identical network on identical shapes, so diffing tensor/two_layer_net_tensor.py against autograd/two_layer_net_autograd.py shows exactly what autograd removes, and diffing that against nn/two_layer_net_module.py shows what nn.Module adds. The custom Function in autograd/two_layer_net_custom_function.py is a clean minimal template: save what backward needs with ctx.save_for_backward, and clone grad_output before masking so the incoming gradient is not mutated in place. The numpy warm-up hand-derives the backward pass, which is the most useful thing here for anyone who has only used frameworks and never wondered what loss.backward() replaces. dynamic_net.py makes define-by-run concrete by picking a random number of reused hidden layers on each forward pass.
The examples have aged out. The prose still recommends reduction='elementwise_mean', which has been superseded by 'mean', and the TensorFlow section is 1.x graph code (tf.placeholder, tf.Session) that does not run as written on TensorFlow 2. Correctness is checked by eye: the scripts print the loss and never assert that it falls, nothing seeds the RNG so runs differ, and you have to read the output to know training worked. The whole thing is one fixed toy problem with random data, so it says nothing about data loading, batching, train versus eval mode, checkpointing, or mixed precision, and GPU use means uncommenting a device line by hand in each file.