// the find
GeekAlexis/FastMOT
High-performance multiple object tracking based on YOLO, Deep SORT, and KLT 🚀
FastMOT is a multi-object tracker for video that combines a YOLO or SSD detector, OSNet re-identification in a Deep SORT-style association step, and KLT optical flow. It runs the detector and feature extractor only every N frames and fills the gaps with KLT, which is what makes real-time speeds on Jetson-class hardware plausible. It suits people building counting or surveillance pipelines on NVIDIA hardware.
- Detection and ReID run every N frames while KLT carries tracks in between, so the expensive network cost is amortized rather than paid per frame. The MOT20 table shows MOTA dropping only about 1.7 points from N=1 to N=5, which is a useful trade-off to see reported.
- Inference runs through TensorRT with asynchronous execution, and the Kalman filter, KLT and data association are Numba-compiled. That is the right place to spend optimization effort on Jetson, where Python overhead would otherwise dominate.
- Camera motion compensation is included, which matters for footage from a moving camera where a static-scene assumption breaks down. Detectors and ReID models can be swapped by subclassing and editing a JSON config instead of forking the tracker.
- The dependency stack is frozen around 2020-2021: Numba 0.48, CuPy 9.2, TensorFlow below 2.0, and an ONNX 1.4.1 pin for conversion. Last push was July 2024. Getting this running on a current CUDA, Python or JetPack will likely take real effort before any tracking happens.
- The 50-150 FPS desktop figure is an expectation, not a measurement. The only benchmarked numbers are MOT17 sequences on a Jetson Xavier NX, so you should expect to benchmark your own hardware.
- The default YOLOv4 weights are trained on CrowdHuman, so the defaults are effectively pedestrian-tuned. Multi-class support is described, but there are no reported numbers for vehicles or mixed classes, and custom ReID means training with fast-reid and converting the weights yourself.
- The install story is Ubuntu-only. The Dockerfile targets Ubuntu 18.04 and 20.04 with old driver minimums, and the README gives no Windows or macOS path. Switching video backends means editing a hard-coded WITH_GSTREAMER flag in videoio.py.