FailSafe SWARM is #1 on CVE-Bench

Trusted by leading technology companies worldwide.

    • NVIDIA
    • Vercel
    • Grab
    • NVIDIA
    • OpenAI
    • Anthropic
    • Coinbase
    • OpenAI
    • NEAR
    • AWS
    • Base
    • NEAR
    • Consensys
    • Ensign
    • OpenEden
    • Consensys

Proof system

Security claims you can check.

Versioned benchmarks, public traces, and coordinated disclosures make our work inspectable.

As of August 2026, FailSafe SWARM holds the highest reported score on CVE-Bench v2.1.0: 62.5% zero-day and 70% one-day at pass@1 (28 of 40 targets), graded by a deterministic oracle with results published under MIT.

Check the evidence

FailSafe's AttackBench, developed with NEAR, ran 624 hostile exchanges between an attacker model and defending AI agents across three runtimes.

Check the evidence

FailSafe has disclosed 240+ vulnerabilities across 101 coordinated reports, including findings at Deutsche Bank, MUFG, Zurich Insurance, and Vercel.

Check the evidence
12 days12 hours

Find and Fix Critical Vulnerabilities That AI Hackers Are Targeting Today

Launch a pentest in minutes and get validated findings with an audit-ready report the same day.

Exploit Verified

Validated Findings and Automated Remediation Guidance

Every issue is proven and reproducible, with clear remediation your team can act on immediately.

MAS
SOC 2
ISO 27001

Pass Your Toughest Audits

Generated reports are explicitly mapped to the controls your auditors care about.

Gartner

By 2028, over 60% of enterprise pen test programs will operate as continuous validation, replacing annual assessments as the primary proof of resilience.

Gartner·How to Implement a Continuous Offensive Security Testing Program, 2026

API & Backend

Web Applications

Continuous penetration testing for APIs, web apps, and cloud infrastructure. Findings mapped to MITRE ATT&CK and OWASP Top 10.

Explore the security frontier

From AI agents to the systems they can reach.

Choose a focused path into FailSafe's ACOST practice, research, and AI security work.

Recognition

Industry Recognition

EDB

Global Founder Programme

Anthropic

Cyber Partner

Meet The Drapers

2025 Winner

Regulation Asia

Best Security 2025

Immunefi

Top 3 Auditor

CoinDesk

Consensus 2024 Winner

Codehawks

Top 3 Auditor

AWS

2026 Partner

CREST Certification

Questions & answers

Frequently asked questions

Quick answers about FailSafe's services, coverage, and engagement process.

FailSafe Security is a cyber frontier lab building Agentic Continuous Offensive Security Testing (ACOST) systems for the AI era. It combines GlassBreak cyber-model R&D, SWARM agentic execution, CTEM workflows, human-led security research, and targeted security assessments for critical systems.

FailSafe offers ACOST and CTEM through SWARM; AI and agent security assessments; penetration testing for applications, APIs, and infrastructure; code assurance; incident response; compliance guidance; and security research.

FailSafe works across AI and agentic systems, web applications, APIs, cloud infrastructure, identity, data, and other critical software systems. Specialized code and financial-system work remains available when the threat model requires it.

Organizations can start a security scan or contact the FailSafe team with their scope, environment, and timeline. The team then recommends an assessment or continuous validation approach suited to the system.