finds.dev← search

// the find

triple-mu/YOLOv8-TensorRT

★ 1,812 · Python · MIT · updated Aug 2026

YOLOv8 using TensorRT accelerate !

A set of Python and C++ tools that take an ultralytics YOLOv8 model through ONNX to a TensorRT engine and run it for detection, segmentation, pose, oriented boxes and classification. It is aimed at people deploying YOLOv8 on NVIDIA hardware, including Jetson and DeepStream, who need more control than the ultralytics runtime gives them.

- TensorRT version branching is confined to one layer (trt_compat) rather than spread through the C++ apps. That matters for a repo that claims to build against TensorRT 8.6 through 11.0.

- The build detects TensorRT and OpenCV versions on its own and switches to class-aware NMS when OpenCV is 4.7 or newer, so the same tree builds across several machines without source edits.

- Python and C++ share engines and label files, and infer.py replaces ten per-task scripts with one entry point and a choice of backends (torch, cudart, pycuda). Comparing a Python run against the C++ binary on the same engine is straightforward.

- The benchmark table states the GPU, CUDA and TensorRT versions and the model size, and it measures host-to-host time including copies, which is the number that matters for a real pipeline.

- There are three export paths with three different engine output layouts: End2End from export-det.py, outputs plus proto from export-seg.py, and raw from ultralytics. The README needs a whole paragraph to say which engine goes with which binary, and a mismatch surfaces at runtime, not at build time.

- Pose, oriented boxes and classification only have the raw ultralytics export, and the README documents no End2End option for them. Their post-processing depends on host code that the README does not describe.

- All benchmark numbers come from one laptop GPU (RTX 3080 Ti Laptop) and only yolov8n. Jetson and DeepStream appear under Deployment with no measurements, though those are probably where most people run this.

- Setup assumes CUDA and TensorRT are already installed system-wide, and TensorRT 8 needs a separate cuDNN 8 to link. The Troubleshooting section is long because the install path is fragile, so expect to spend time on library paths before the first inference runs.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →