finds.dev← search

// the find

weicj/vLLM-2080Ti-Definitive

★ 1,109 · Python · Apache-2.0 · updated Oct 2026

The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with maximum 200+ tok/s single-request decode with support of FP8 weight ( Join Discord :https://discord.gg/VFqVVySdMS )

A fork of vLLM with patches targeting old Turing-architecture GPUs (2080 Ti, Tesla T10, Quadro RTX 6000/8000) for local LLM serving. It's aimed at people who bought cheap 22GB+NVLink 2080 Ti pairs or picked up decommissioned T10 cards and want to run Qwen 27B/35B class models instead of buying a 3090/4090.

The launcher and profile system (profiles/<hardware>/<model>/<weight>/<route>.env) is genuinely practical — it separates GPU topology, quant scheme, and KV-cache dtype into reusable configs instead of a pile of shell flags. It also covers real TP=2/TP=4 multi-GPU layouts over PCIe, not just the NVLink pair, which is the part that usually breaks on older hardware. The hardware Q&A section is honest about limits that most forks gloss over: P2P vs NVLink matters, mixed 11GB/22GB cards won't work for the documented routes, and thermal throttling can masquerade as a software regression.

The headline comparison table (2x2080Ti vs 3090 Ti) stacks raw SM/tensor-core counts across two different Nvidia generations without adjusting for the fact that Ampere tensor cores do roughly 2x the work per core — so '3.24x more tensor cores' implies a performance edge the hardware doesn't actually have. The directory tree is the full stock vLLM tree (CPU/ARM/RISC-V kernels, ROCm Kimi-K3 kernels, every upstream benchmark) rather than an isolated patch set, which doesn't match the 'hardware-focused fork' framing and makes it hard to tell what was actually changed for SM75 versus what's just inherited upstream bloat. The benchmark numbers (220 tok/s) are explicitly synthetic single-request best cases with 'high speculative-hit rates' and the README admits real workloads won't hit them, which undercuts the headline stat. Single maintainer, Discord-driven support model, and no CI/test badges visible — for a build this fiddly (CUDA 13 + specific PyTorch + kernel version alignment), that's a real risk when upstream vLLM moves fast and this fork has to keep rebasing.

View on GitHub →

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →