finds.dev← search

// the find

BitterSecurity/Decepticon

★ 5,581 · Python · Apache-2.0 · updated Sep 2026

Autonomous Hacking Agent for Red Team

Decepticon is a LangGraph-orchestrated red-team agent that runs a full kill chain — recon through C2 — inside an isolated Docker sandbox, with 16 specialist agents and generated engagement paperwork (RoE, ConOps, OPPLAN) before it fires a single command. Built for pentesters and red teamers who want an LLM to drive real offensive tooling rather than just summarize an nmap scan.

Commands run in persistent tmux sessions with prompt detection, so interactive tools like msfconsole, sliver-client, and evil-winrm actually work end-to-end instead of failing on the first prompt. The two-network split (management plane vs. sandbox/C2/target network) is a real architectural boundary, not just a diagram. It ships reproducible benchmark traces (LangSmith links, per-challenge JSON/MD reports) against XBOW rather than just claiming a pass rate.

The XBOW validation set it's scored against is maintained by the same org (PurpleAILAB), so the 98% headline number has no independent third party behind it. Full stack is LiteLLM + Postgres + Neo4j + sandbox + on-demand BloodHound CE/Sliver C2/Ghidra MCP — a lot of moving Docker parts to keep patched and a lot of attack surface for something that's itself offensive tooling. Handing an LLM live control of a C2 framework and lateral-movement primitives is a real misuse vector; the README's guardrail is a 'safety gate' plus a disclaimer, not a technical control anyone outside the project has audited.

View on GitHub → Homepage ↗

// want more like this?

We dig through GitHub every week and send a few repos picked for what you actually care about — each with an honest take like this one.

Get finds in your inbox → Search again →