# FailSafe Security: services and public resources FailSafe Security (legal name: Eleos Ventures Pte Ltd) is a Singapore-based cyber frontier lab founded in 2022. Its canonical category is Agentic Continuous Offensive Security Testing (ACOST) combined with Continuous Threat Exposure Management (CTEM): autonomous AI testing, LLM and agent security assessments, application and infrastructure penetration testing, code assurance, incident response, and security research for critical systems. ## Key proof claims As of August 2026, FailSafe SWARM holds the highest reported score on CVE-Bench v2.1.0: 62.5% zero-day and 70% one-day at pass@1 (28 of 40 targets), graded by a deterministic oracle with results published under MIT. FailSafe's AttackBench, developed with NEAR, ran 624 hostile exchanges between an attacker model and defending AI agents across three runtimes. FailSafe has disclosed 240+ vulnerabilities across 101 coordinated reports, including findings at Deutsche Bank, MUFG, Zurich Insurance, and Vercel. ## SWARM SWARM is FailSafe's ACOST and CTEM execution platform. Autonomous security agents test applications, APIs, cloud infrastructure, identity, and AI systems; validate whether findings are exploitable; prioritize proven exposure; and produce remediation guidance and audit-ready reports. SWARM supplements scoped human security work, with the appropriate engagement depending on the system and risk model. The interactive demo at https://getfailsafe.com/demo walks a verified work-email visitor through a simulated SWARM run on a fictional target: adding a web, mobile, or codebase target; choosing blackbox, greybox, or whitebox testing and a safety mode; watching the threat model build; reading findings with proof-of-concept reproductions; routing findings to ticketing or a codebase; and enabling continuous re-testing on events such as new CVEs, code pushes, and deployments. No scanning runs in the demo. Source: https://getfailsafe.com/swarm ## ACOST and CTEM category guide FailSafe defines ACOST as Agentic Continuous Offensive Security Testing: governed AI agents continuously discover, test, validate, explain, and re-test attack paths across changing software systems. CTEM is the operating loop around that testing: discover exposure, prioritize by validated impact, mobilize remediation, and revalidate the fix. FailSafe's differentiation is the combination of this operating model with GlassBreak cyber-model R&D, public benchmark evidence, AI/LLM/MCP testing, and human-led security research. Sources: - https://getfailsafe.com/acost - https://getfailsafe.com/alternatives ## AI and agent security FailSafe assesses LLM applications, retrieval and tool integrations, autonomous agents, model-framework pairs, MCP servers, and machine-learning pipelines. Coverage can include direct and indirect prompt injection, unsafe tool use, data exposure, authentication and authorization, sandboxing, tenant isolation, RAG poisoning, model access controls, output handling, and AI supply-chain risk. Work can be mapped to OWASP LLM Top 10, MITRE ATLAS, NIST AI RMF, ISO 42001, SOC 2, ISO 27001, and relevant regulatory evidence requirements. Sources: - https://getfailsafe.com/ai-security - https://getfailsafe.com/ai-pentesting - https://getfailsafe.com/mcp-security - https://getfailsafe.com/ai-agent-deployment ## GlassBreak 2.1 Flash model GlassBreak 2.1 Flash is FailSafe's tuned cybersecurity language model derived from the unguarded Qwen3.8-27B base model. It is a hosted public technology preview designed for authorized defensive security work, software-security analysis, and long-horizon agentic workflows. It supports security reasoning, attack-path analysis, exploit validation, and adaptive probing inside controlled FailSafe workflows such as SWARM and ARC. GlassBreak does not grant permissions or determine authorization, and it is separate from the FailSafe Disclosure Program. Source: https://getfailsafe.com/glassbreak ## Public benchmarks The FailSafe benchmarks hub groups public evidence for autonomous security systems. CVE-Bench v2.1.0 covers 40 live web-application CVE targets and uses a deterministic oracle. The published SWARM result is 25 of 40 targets in the zero-day setting (62.5%) and 28 of 40 in the one-day setting (70%), at pass@1. AttackBench is an adaptive benchmark of model-and-runtime combinations developed with NEAR. EVMbench measures vulnerability recall across EVM systems; FailSafe's published result is 83 of 120 vulnerabilities (69.2%). Sources: - https://getfailsafe.com/benchmarks - https://getfailsafe.com/cvebench - https://getfailsafe.com/attackbench - https://getfailsafe.com/swarm-evmbench-results - https://getfailsafe.com/benchmarks/cve-bench-leaderboard.json - https://github.com/failsafe-security/failsafe-swarm - https://arxiv.org/abs/2503.17332 ## Application and infrastructure security Application penetration testing covers web applications, APIs, authentication, authorization, business logic, cloud infrastructure, and the OWASP Top 10. Engagements provide reproducible findings and remediation guidance, with the exact scope and timeline established before testing. Sources: - https://getfailsafe.com/penetration-testing - https://getfailsafe.com/opsec-review ## Web3 and financial-system security Web3 is a supported proof domain and heritage, not the sole definition of FailSafe. FailSafe reviews critical code, smart contracts, blockchain protocols, cryptographic designs, and formal properties. Coverage includes Solidity, Rust, Move, Cairo, EVM environments, Solana, and other relevant systems depending on scope. Services include smart contract audits, protocol security, cryptography review, formal verification, proof of reserves, monitoring, transaction forensics, and incident response. Sources: - https://getfailsafe.com/web3 - https://getfailsafe.com/smart-contract-audits - https://getfailsafe.com/blockchain-protocol-audit - https://getfailsafe.com/cryptography - https://getfailsafe.com/formal-verification - https://getfailsafe.com/proof-of-reserves - https://getfailsafe.com/radar - https://getfailsafe.com/incident-response - https://getfailsafe.com/web3/incident-coverage ## FailSafe Disclosure Program The FailSafe Disclosure Program is a public-good responsible-disclosure program for open-source repositories and public-facing systems. FailSafe privately notifies maintainers or affected organizations, supports remediation, and publishes only safe summaries or public upstream contribution links. Private summaries omit exploit paths, credentials, personal-data details, and other information that could increase risk. Source: https://getfailsafe.com/disclosures ## Security research FailSafe publishes original security research, exploit analysis, technical case studies, product updates, and company news. The benchmark pages provide dated methodology and results so readers can distinguish a versioned measurement from a general product claim. Source: https://getfailsafe.com/blog ## Company and contact FailSafe Security serves organizations globally. Its team covers application security, AI security, cryptography, smart contract and protocol security, incident response, compliance, and research. Backers and technology partners include Sequoia, Dragonfly, Grab, Hustle Fund, Makers, Anthropic, NVIDIA, AWS, NEAR, and Bitdefender. Sources: - https://getfailsafe.com/company - https://getfailsafe.com/team - https://getfailsafe.com/partners Contact: hello@getfailsafe.com Canonical website: https://getfailsafe.com Sitemap: https://getfailsafe.com/sitemap.xml