// the find
jacobgil/pytorch-grad-cam
Advanced AI Explainability for computer vision. Support for CNNs, Vision Transformers, Classification, Object detection, Segmentation, Image similarity and more.
pytorch-grad-cam is a Python library that implements a very large set of class-activation-map and pixel-attribution methods (GradCAM, ScoreCAM, AblationCAM, EigenCAM, LayerCAM, FullGrad, ShapleyCAM, and a dozen more) for explaining what CNNs and Vision Transformers are looking at. It covers classification, object detection, segmentation, and embedding/CLIP similarity, so it's aimed at ML engineers and researchers who need to visualize or debug model attention rather than just read a paper about it.
The breadth is the main draw: one consistent API gives you access to 18+ CAM variants, so swapping GradCAM for ScoreCAM or EigenCAM is a one-line change instead of re-deriving gradient hooks yourself. It isn't locked to CNN classifiers either — the reshape_transform and model_target abstractions let it handle Vision Transformers, object detection, segmentation, and CLIP text-image similarity, which is unusual for a CAM library. It also ships real quantitative faithfulness metrics (ROAD, ARCC) instead of just producing heatmaps, so you can actually argue one explanation method is better than another rather than eyeballing it. The context-manager API properly releases forward/backward hooks, and there's a dedicated test for hook/context release on CPU and CUDA, which suggests the maintainer has been burned by hook leaks before and fixed it properly.
There are no type hints visible anywhere in the core constructor or target/reshape-transform signatures, which hurts for a library this configurable — you'll be reading source to know what shape a reshape_transform is expected to return. Documentation beyond the README is a jupyter-book of tutorials rather than a generated API reference, so understanding the exact semantics of the more niche methods (FEM, RefineCAM, SESS, ShapleyCAM) means reading the implementation directly. AblationCAM and ScoreCAM need many forward passes per image even with batching, and there are no concrete runtime numbers given anywhere, so you discover the actual cost on your own hardware. The test suite is also thin relative to the method count — roughly ten test files cover 18+ CAM variants, so correctness of the newer or more exotic methods is harder to verify independently of trusting the author.