finds.dev← search

// the find

soumith/convnet-benchmarks

★ 2,688 · Python · MIT · updated Jun 2017

Easy benchmarking of all publicly accessible implementations of convnets

A benchmark suite comparing forward/backward pass speed of convolution kernels across Caffe, Torch, TensorFlow, Theano, Chainer, cuDNN, and a half-dozen other frameworks circa 2014-2015. Useful for anyone doing historical research into GPU convolution performance or curious how the deep learning framework landscape looked before cuDNN standardized things.

Actually reproducible: each framework directory ships the exact script, prototxt, or config used to generate its numbers, plus raw output logs, not just a table someone typed up. The layer-wise breakdown (L1-L5) isolates where time goes for kernels of different sizes and channel counts, which is more informative than a single end-to-end number. Comparing native Caffe/Torch conv against fbfft and cuDNN R2/R4 side by side captures a real inflection point in how convolutions were computed.

Last commit is from mid-2017 and the hardware is a Titan X with CUDA-era cuDNN R2/R4 — none of this reflects Ampere/Hopper, Tensor Cores, or any cuDNN version from the last eight years, so the numbers are historical curiosities, not a buying guide. Half the frameworks benchmarked (Chainer, Nervana neon, cuda-convnet2, DeepCL, Caffe-CLGreenTea) are dead projects, so most of the comparison table is now benchmarking vaporware against vaporware. No CI, no way to re-run this against a current GPU or current framework versions — it's a frozen snapshot with no path to becoming a living benchmark again.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →