// the find
lllyasviel/stable-diffusion-webui-forge
A fork of AUTOMATIC1111's Stable Diffusion WebUI focused on faster inference, lower VRAM usage, and quicker support for new model architectures (Flux, SD3.5, Chroma, Kolors). It's for people already running local Stable Diffusion who've hit memory or speed ceilings on the original webui, especially anyone trying to run Flux on consumer GPUs via quantized weights.
Ships a real GPU memory management system plus native support for quantized Flux (BNB NF4, GGUF Q8/Q5/Q4) with LoRA support and a GPU-weight/offload slider to trade VRAM for speed. The UnetPatcher pattern (shown in the FreeU example) lets you hook into the sampling pipeline from a single extension file instead of forking core sampler code. It stays compatible with the existing AUTOMATIC1111 extension ecosystem (ControlNet, IP-Adapter, InstantID, Adetailer) so users don't lose their extension stack switching over.
The README's own 'Forge Status' table is a manually-updated, self-admitted list of broken features (Surface pen pressure, OFT LoRAs, ControlNet Union/Flux) with no fix timeline — you're told upfront some things just don't work. It only syncs with upstream AUTOMATIC1111 every 90 days, so base-webui fixes lag behind. Install is Windows/CUDA-first via versioned .7z packages tied to specific CUDA/Torch combos; anything else falls back to a manual git-clone path under 'Advanced Install.' Documentation is mostly a list of links into GitHub Discussions threads (including basics like fixing 'Connection errored out'), so troubleshooting means searching forum posts rather than reading docs.