// the find
BitterSecurity/Decepticon
Autonomous Hacking Agent for Red Team
Decepticon is a LangGraph-orchestrated red-team agent that runs a full kill chain — recon through C2 — inside an isolated Docker sandbox, with 16 specialist agents and generated engagement paperwork (RoE, ConOps, OPPLAN) before it fires a single command. Built for pentesters and red teamers who want an LLM to drive real offensive tooling rather than just summarize an nmap scan.
Commands run in persistent tmux sessions with prompt detection, so interactive tools like msfconsole, sliver-client, and evil-winrm actually work end-to-end instead of failing on the first prompt. The two-network split (management plane vs. sandbox/C2/target network) is a real architectural boundary, not just a diagram. It ships reproducible benchmark traces (LangSmith links, per-challenge JSON/MD reports) against XBOW rather than just claiming a pass rate.
The XBOW validation set it's scored against is maintained by the same org (PurpleAILAB), so the 98% headline number has no independent third party behind it. Full stack is LiteLLM + Postgres + Neo4j + sandbox + on-demand BloodHound CE/Sliver C2/Ghidra MCP — a lot of moving Docker parts to keep patched and a lot of attack surface for something that's itself offensive tooling. Handing an LLM live control of a C2 framework and lateral-movement primitives is a real misuse vector; the README's guardrail is a 'safety gate' plus a disclaimer, not a technical control anyone outside the project has audited.