FailSafe SWARM is #1 on CVE-Bench

Agentic Continuous Offensive Security Testing (ACOST) / Public proof

Measure the attacker. Publish the evidence.

FailSafe is a cyber frontier lab benchmarking autonomous security systems against real vulnerabilities and real agent runtimes. Results are versioned, oracle-graded where applicable, and published for inspection.

Public results

A benchmark is only useful when others can check it.

CVE-Bench v2.1.0

Autonomous black-box exploitation

62.5% zero-day · 70% one-day

FailSafe SWARM exploited 25 of 40 targets in the zero-day setting and 28 of 40 in the one-day setting at pass@1, graded by a deterministic oracle.

Published report and traces

AttackBench × NEAR

Model + runtime safety

624 hostile exchanges

An adaptive attacker tested twelve model-framework combinations across three agent runtimes, showing that safety is a property of the model and runtime together.

Read the research

EVMbench

Autonomous vulnerability discovery

69.2% recall · 83/120

SWARM's published EVMbench result measures recall across 120 vulnerabilities and compares the result with other capable models.

Read the results

Proof system

Security claims you can check.

Versioned benchmarks, public traces, and coordinated disclosures make our work inspectable.

As of August 2026, FailSafe SWARM holds the highest reported score on CVE-Bench v2.1.0: 62.5% zero-day and 70% one-day at pass@1 (28 of 40 targets), graded by a deterministic oracle with results published under MIT.

Check the evidence

FailSafe's AttackBench, developed with NEAR, ran 624 hostile exchanges between an attacker model and defending AI agents across three runtimes.

Check the evidence

FailSafe has disclosed 240+ vulnerabilities across 101 coordinated reports, including findings at Deutsche Bank, MUFG, Zurich Insurance, and Vercel.

Check the evidence

Questions & answers

Frequently asked questions

How FailSafe measures autonomous security systems and publishes the evidence.

FailSafe publishes SWARM results on CVE-Bench, the AttackBench adaptive benchmark developed with NEAR, and EVMbench results for autonomous vulnerability discovery in EVM systems.

On CVE-Bench v2.1.0 at pass@1, FailSafe SWARM reported 62.5% in the zero-day setting and 70% in the one-day setting across 40 targets. The result is versioned, oracle-graded, and published with per-target evidence.

The benchmark methods, graded results, and evidence are public. Because autonomous systems are stochastic, the published repositories describe the method and evidence rather than guaranteeing an identical score on every future run.

Benchmarks turn broad claims about AI security into measurable, comparable tasks. They help security teams evaluate an agent's ability to discover and validate vulnerabilities under a defined threat model.