// the find
WebFuzzing/EvoMaster
The first open-source AI-driven tool for automatically generating system-level test cases (also known as fuzzing) for web/enterprise applications. Currently targeting whitebox and blackbox testing of Web APIs, like REST, GraphQL and RPC (e.g., gRPC and Thrift).
An evolutionary-algorithm-based fuzzer for REST, GraphQL, and RPC (gRPC/Thrift) APIs that generates system-level test cases and regression suites, not just crash reports. Aimed at teams testing JVM-based APIs who want more than schema-driven black-box fuzzing — the white-box mode does bytecode analysis to guide test generation toward actual code paths.
White-box mode analyzes bytecode (taint analysis, testability transformations) to generate tests that hit real branches, which is a genuine step up from schema-only fuzzers like Schemathesis or Dredd — and two independent academic studies back up the coverage/fault-finding claims against competitors. It understands SQL and MongoDB traffic during white-box runs and can seed the database directly, with that setup baked into the generated tests. Output isn't just a bug report — it's runnable JUnit/pytest/Jest suites that start and stop the app themselves, so they drop into an existing CI pipeline as regression tests. 30+ built-in fault oracles (500s, schema mismatches, BOLA, SQLi) is a lot of detection logic you don't have to write yourself.
White-box mode — the mode that actually produces good results — requires hand-writing a driver for your specific application; black-box mode is the 5-minute path but the README admits results are worse. Bytecode instrumentation gets harder with every JDK release past 8, and the docs have a whole page dedicated to workarounds for anything above it. RPC fuzzing has no schema support at all: you write a driver using the service's own client library, so there's no equivalent of pointing it at an OpenAPI file like you can for REST. It also can't mock external service calls yet, so any SUT that reaches out to other APIs during a test run is still an open problem, and getting results worth trusting means running it for hours, not minutes — not a tool you fit into a quick PR check.