// the find
elder-plinius/OBLITERATUS
OBLITERATE THE CHAINS THAT BIND YOU
OBLITERATUS is a toolkit for abliteration — using activation-space analysis to locate and surgically remove refusal behavior from open-weight LLMs, without fine-tuning. It ships as a CLI, Python API, and a zero-code Colab/HuggingFace Spaces UI, and is aimed at anyone who wants an uncensored local model, dressed in the language of alignment research.
It implements a real, citable technique (Arditi et al.'s refusal-direction work) with actual engineering behind it: multiple extraction methods (mean-difference, PCA, whitened SVD), 15 separate analysis modules, and an 'informed' pipeline that uses those analyses to auto-configure the removal step instead of one blunt pass. It also handles production realities most scripts like this skip — MoE fused-expert tensors, accelerate-offloaded meta-tensor layers, FP8/NVFP4 dequantization — and offers a reversible steering-vector path alongside the permanent weight-projection one, so you're not locked into a destructive edit just to test the effect.
The 'for researchers, not bad actors' framing is an honor system with no technical gate — the one-click Colab/Spaces path removes exactly the friction that would separate a red-teamer from someone who just wants an unfiltered chatbot. Telemetry is on by default on Spaces and feeds the maintainers' own research dataset, which is easy to miss if you're just trying the tool. The bundled model shortlist already includes pre-jailbroken checkpoints (Dolphin, Hermes, WhiteRabbitNeo) for comparison, which undercuts the stated research purpose further. The multi-GPU/multi-host story is also weaker than the README's length suggests — it's single-process device placement plus a preflight checker, not real distributed checkpoint support, and that caveat is buried deep in the docs.