Prove
Is the suspected vulnerability exploitable in this environment, or can it be defensibly refuted?
FailSafe category / ACOST
ACOST is FailSafe's operating category for governed AI agents that continuously discover, test, validate, explain, and re-test attack paths across changing software systems.
FailSafe operationalizes CTEM by continuously discovering, validating, prioritizing, and re-testing exploitable exposure. GlassBreak cyber-model R&D, SWARM agentic execution, public benchmark evidence, and human-led security research work together as one offensive security system.
The acceptance criterion
ACOST is not an LLM summary layered on top of a scanner. The acceptance criterion is a reproducible exploit or a defensible refutation, connected to remediation guidance and post-fix proof.
Is the suspected vulnerability exploitable in this environment, or can it be defensibly refuted?
What engineering change removes the demonstrated risk while preserving intended behavior?
Did remediation close the attack path, with reproducible evidence for security and risk owners?
Attack surface
Evidence contract
ACOST becomes useful when the security team agrees in advance what a valid outcome looks like. These criteria keep agentic testing measurable, bounded, and useful to engineering and risk owners.
A security-relevant effect is demonstrated in the authorized target environment, not inferred from a version string or pattern match.
The trace preserves the hypothesis, actions, responses, target context, and grader or verification outcome.
Authorization, target boundaries, tool permissions, rate limits, and high-impact approvals are explicit before execution.
The same attack path or a defined regression test is re-run after remediation, with the result recorded for the risk owner.
Cyber frontier lab
FailSafe develops GlassBreak cyber models, runs them through controlled agentic harnesses, and publishes benchmark evidence so security teams can inspect the method instead of relying on a broad capability claim.
Agentic Continuous Offensive Security Testing (ACOST) is the product category. GlassBreak is the model R&D layer. SWARM is the agentic execution layer. Human researchers provide scoping, judgment, verification, and escalation for high-impact work.
Read the public benchmark evidence or compare FailSafe with other platforms evaluating the same category.
Proof system
Versioned benchmarks, public traces, and coordinated disclosures make our work inspectable.
As of August 2026, FailSafe SWARM holds the highest reported score on CVE-Bench v2.1.0: 62.5% zero-day and 70% one-day at pass@1 (28 of 40 targets), graded by a deterministic oracle with results published under MIT.
Check the evidenceFailSafe's AttackBench, developed with NEAR, ran 624 hostile exchanges between an attacker model and defending AI agents across three runtimes.
Check the evidenceFailSafe has disclosed 240+ vulnerabilities across 101 coordinated reports, including findings at Deutsche Bank, MUFG, Zurich Insurance, and Vercel.
Check the evidenceQuestions & answers
What ACOST means and how FailSafe combines agentic offensive security testing with CTEM.
ACOST means Agentic Continuous Offensive Security Testing. It describes governed AI agents that continuously discover, test, validate, explain, and re-test attack paths across changing software systems.
Vulnerability scanning usually identifies suspected weaknesses or matches known patterns. ACOST uses agentic exploration and attack-path validation to establish whether a weakness is exploitable in context, then connects evidence to remediation and retesting.
ACOST provides the continuous offensive testing engine. CTEM provides the operating loop around it: discover exposure, prioritize by validated impact, mobilize remediation, and revalidate the fix over time.
FailSafe tests applications, APIs, cloud and identity controls, LLM applications, AI agents, MCP servers, tools, data flows, and other critical software systems. Scope and permissions are defined before testing begins.
Evidence is a recorded, scope-compliant interaction that demonstrates a security-relevant effect or defensibly refutes one. A useful result preserves the target context, actions, responses, outcome, and remediation state so an independent reviewer can understand and repeat the conclusion.