// the find
soumith/cudnn.torch
Torch-7 FFI bindings for NVIDIA CuDNN
FFI bindings that expose NVIDIA cuDNN's convolution, pooling, batchnorm, and RNN kernels as drop-in replacements for Torch7's nn modules. It's for people maintaining old Torch7 codebases who want GPU-accelerated primitives without hand-writing CUDA, not for anyone starting a new project today.
cudnn.convert swaps metatables in place with no memory copy, so converting an existing nn graph to cudnn (or back) is nearly free. Modules are unit-tested against the nn reference implementations, not just assumed correct. Coverage is broad — 2D/3D conv and pooling, LSTM/GRU/BLSTM, batchnorm, several softmax variants — rather than a thin wrapper around one op. The benchmark/fastest/verbose flags expose the actual perf-vs-memory tradeoff instead of hiding an opinionated default.
Torch7 is dead and so is this repo — last push was 2018, pinned to cuDNN R5 EA and CUDA 7.0, which won't build against anything from the last several CUDA generations. Install is manual: find libcudnn.so, put it on LD_LIBRARY_PATH or copy it into the CUDA dir yourself, no packaging beyond a rockspec. No half/mixed-precision support and nothing from cuDNN R6 onward. The 'no backward pass with batchnorm in eval mode' caveat is a real correctness footgun and it's buried as a single backticked sentence in the README rather than surfaced anywhere in code.