// the find
jcjohnson/torch-rnn
Efficient, reusable RNNs and LSTMs for torch
torch-rnn provides RNN and LSTM modules for Torch7, plus a character-level language model that reimplements Karpathy's char-rnn for speed. It is for people with an existing Lua/Torch codebase who want the recurrent modules without extra dependencies. Most of the field has moved to PyTorch.
The modules in LSTM.lua and VanillaRNN.lua depend only on torch and nn, so you can lift them into another project without the training pipeline. The benchmark section states its configuration (tiny-shakespeare, batch 50, sequence length 50, no dropout, first 100 iterations), so the 1.9x and 7x figures can be checked against a stated setup. The tests include a gradient checker (util/gradcheck.lua) and a comparison against Zaremba's reference LSTM, which is the right way to validate hand-written recurrent code.
Setup means Torch7, luarocks packages, a deepmind fork of torch-hdf5, and Python 2.7 headers, and the repo has not been pushed since June 2022. Expect installation to take longer than a training run. The preprocessing script is Python 2.7 and writes HDF5 plus JSON, and the README's own TODO admits those dependencies should go. The model is character-level only, with no attention or tokenization, so it teaches RNN mechanics well but is not a base for anything current. The speedup numbers come from a Titan X and an i7-4790k, both several GPU generations old, compared against char-rnn, which tells you little about cuDNN LSTMs in PyTorch today.