FailSafe SWARM is #1 on CVE-Bench

Agentic Continuous Offensive Security Testing (ACOST)

Secure the systems that can act for you.

FailSafe is a cyber frontier lab building Agentic Continuous Offensive Security Testing (ACOST) systems for the AI era. Test AI products at the layer where risk compounds: the model, the runtime, the tools, the data, and the production environment around them.

FailSafe operationalizes CTEM by continuously discovering, validating, prioritizing, and re-testing exploitable exposure.

What we mean

AI security is bigger than the model.

Agentic AI security is the practice of testing and hardening AI systems that can retrieve information, call tools, change data, or take actions.

That means testing the instructions the model receives, the data it can see, the tools it can call, and the controls that govern every action. A capable model can still become unsafe when a framework renders untrusted data as instructions or grants a tool more access than the workflow requires.

Our ACOST work combines focused LLM assessments with continuous offensive validation across the rest of the application and infrastructure stack, giving CTEM teams evidence they can act on and re-test.

Coverage

Map the full AI attack surface.

Agents & runtimes

Red-team planning, memory, permissions, multi-agent orchestration, and tool-call behavior.

LLM applications

Prompt injection, jailbreaks, output handling, access controls, and cross-user data exposure.

RAG & data

Retrieval poisoning, context contamination, sensitive-data leakage, and untrusted content.

Governance & evidence

Risk-based testing mapped to AI security frameworks and audit-ready remediation evidence.

Skills & orchestration

Agent identity, memory, multi-agent messages, recursive tool use, and workflow-level failure modes.

Proof system

Security claims you can check.

Versioned benchmarks, public traces, and coordinated disclosures make our work inspectable.

As of August 2026, FailSafe SWARM holds the highest reported score on CVE-Bench v2.1.0: 62.5% zero-day and 70% one-day at pass@1 (28 of 40 targets), graded by a deterministic oracle with results published under MIT.

Check the evidence

FailSafe's AttackBench, developed with NEAR, ran 624 hostile exchanges between an attacker model and defending AI agents across three runtimes.

Check the evidence

FailSafe has disclosed 240+ vulnerabilities across 101 coordinated reports, including findings at Deutsche Bank, MUFG, Zurich Insurance, and Vercel.

Check the evidence

Questions & answers

Frequently asked questions

Direct answers about securing LLM applications, agents, and the systems around them.

Agentic AI security is the practice of testing and hardening AI systems that can retrieve information, call tools, change data, or take actions. It covers the model and the runtime around it, including prompts, permissions, tools, memory, data, integrations, and deployment controls.

FailSafe tests prompt injection, indirect prompt injection, unsafe tool use, data leakage, RAG poisoning, tenant isolation, authentication and authorization, output handling, model access controls, and the security of connected APIs and infrastructure.

AI security assessments examine the model, agent, and integration-specific attack surface. SWARM adds continuous validation across the surrounding application, API, cloud, and infrastructure attack surface so fixes can be re-tested as the system changes.

Engagements can be mapped to the OWASP LLM Top 10, MITRE ATLAS, NIST AI RMF, SOC 2, ISO 27001, ISO 42001, EU AI Act evidence needs, and the controls that are relevant to the system's risk model.

Agentic testing should cover the full action loop: untrusted prompts and retrieved content, planning and memory, tool descriptions and arguments, agent identity, least-privilege permissions, multi-agent messages, recursive tool use, secrets, data exfiltration, supply-chain components, and human approval for high-impact actions.