finds.dev← search

// the find

jcjohnson/densecap

★ 1,598 · Jupyter Notebook · MIT · updated Jul 2018

Dense image captioning in Torch

DenseCap finds regions in an image and writes a short caption for each one. It uses a fully convolutional network trained end-to-end on Visual Genome, and this repo is the reference Torch implementation of the CVPR 2016 paper. It is most useful to researchers reproducing the work or studying dense captioning architectures, not to people who want a production captioning tool.

- The localization and captioning share one convolutional backbone, and the localization layer (LocalizationLayer.lua, BilinearRoiPooling, BatchBilinearSamplerBHWD) is trained with the captioning loss. The pieces map cleanly onto the paper's architecture.

- The code is split along the paper's seams. BoxSampler, BoxRegressionCriterion, LanguageModel.lua and the custom criteria each live in their own module with a matching test file under test/, so you can read one piece without loading the rest.

- A pretrained model ships with a download script and a run_model.lua entry point, so you can see output on a single image before committing to the 1.2 GB download or any training setup.

- The vis/ directory has a browser viewer that renders boxes and captions from the output JSON. That is a useful debugging tool for localization quality, which is hard to judge from captions alone.

- It is built on Torch7 and the luarocks ecosystem of that period. Torch7 has been unmaintained for years, and two dependencies (stnbhwd and torch-rnn) are installed from raw GitHub rockspec URLs. Expect the install to be the hardest part, and expect to pin old commits by hand.

- The last push was July 2018. The README still assumes Python 2.7 for evaluation, uses SimpleHTTPServer for the viewer, and links to a Stanford course page for the webcam client. Treat this as a reference snapshot, not a maintained library.

- No dependency versions are pinned anywhere in the README. Reproducing the paper's numbers means matching Torch, cuDNN and nngraph versions from around 2016 to 2018, and nothing in the repo tells you which combination works.

- Training needs the full Visual Genome images, a preprocessing pass that writes one HDF5 file, and realistically a CUDA GPU. Evaluation also needs Java and METEOR. Beyond running inference on a few images, the setup cost is high enough that most people will not get as far as training.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →