// the find
jcjohnson/cnn-vis
Use CNNs to generate images
cnn-vis is a single-script Python tool that generates images by maximizing the activations of a CNN layer, following the Inceptionism post Google published in June 2015. It sits on top of Caffe and the GoogLeNet weights, so it is aimed at researchers or hobbyists who already have a working Caffe install and want to experiment with the visualizations.
The README spells out the layer-amplification objective precisely enough to reimplement: the gradient is -l1_weight * abs(a) - l2_weight * clip(a, -g, g), with the defaults and the reasoning for the thresholded square. The multiscale scheme is implemented as overlapping 224x224 tiles with interleaved updates and upsampling between sizes, which is the part that gives the fractal look without needing a larger network. The regularizers are separated and individually tunable (a p-norm pulling toward the initial image, an auxiliary high-exponent p-norm for saturated pixels, and a TV term with a beta exponent and a schedule), so it works as a research tool rather than a black box. The repo ships 13 example scripts with their outputs, so you can reproduce the gallery before changing anything.
Getting it running is the main obstacle: it needs Caffe built from source with its Python bindings, plus a hand-written .pth file pointing a virtualenv at them, and Caffe itself is effectively unmaintained now. The code targets Python 2.7 and the README's paths assume that, so a modern Python needs a port before it will run. The last push was July 2015, and there are no tests, no CI, and no pinned requirements beyond a requirements.txt. Most of the look comes from the tuning of about thirty flags, and the README gives defaults but little guidance on which knobs to turn for a given kind of image, so expect a lot of trial and error.