Agentic Continuous Offensive Security Testing (ACOST)
Secure the systems that can act for you.
FailSafe is a cyber frontier lab building Agentic Continuous Offensive Security Testing (ACOST) systems for the AI era. Test AI products at the layer where risk compounds: the model, the runtime, the tools, the data, and the production environment around them.
FailSafe operationalizes CTEM by continuously discovering, validating, prioritizing, and re-testing exploitable exposure.
What we mean
AI security is bigger than the model.
Agentic AI security is the practice of testing and hardening AI systems that can retrieve information, call tools, change data, or take actions.
That means testing the instructions the model receives, the data it can see, the tools it can call, and the controls that govern every action. A capable model can still become unsafe when a framework renders untrusted data as instructions or grants a tool more access than the workflow requires.
Our ACOST work combines focused LLM assessments with continuous offensive validation across the rest of the application and infrastructure stack, giving CTEM teams evidence they can act on and re-test.
Coverage
Map the full AI attack surface.
Agents & runtimes
Red-team planning, memory, permissions, multi-agent orchestration, and tool-call behavior.
LLM applications
Prompt injection, jailbreaks, output handling, access controls, and cross-user data exposure.
RAG & data
Retrieval poisoning, context contamination, sensitive-data leakage, and untrusted content.
Governance & evidence
Risk-based testing mapped to AI security frameworks and audit-ready remediation evidence.
Skills & orchestration
Agent identity, memory, multi-agent messages, recursive tool use, and workflow-level failure modes.
Proof system
Security claims you can check.
Versioned benchmarks, public traces, and coordinated disclosures make our work inspectable.
As of August 2026, FailSafe SWARM holds the highest reported score on CVE-Bench v2.1.0: 62.5% zero-day and 70% one-day at pass@1 (28 of 40 targets), graded by a deterministic oracle with results published under MIT.
Check the evidenceFailSafe's AttackBench, developed with NEAR, ran 624 hostile exchanges between an attacker model and defending AI agents across three runtimes.
Check the evidenceFailSafe has disclosed 240+ vulnerabilities across 101 coordinated reports, including findings at Deutsche Bank, MUFG, Zurich Insurance, and Vercel.
Check the evidencePublic frameworks
Ground the assessment in shared language.
AI penetration testing
Test LLM apps, models, agents, RAG, and AI integrations against real attack classes.
ExploreMCP security
Assess Model Context Protocol servers, tool permissions, trust boundaries, and agent-tool abuse.
ExploreSWARM continuous validation
Keep testing the web, API, cloud, and AI surfaces as your product changes.
ExploreQuestions & answers
Frequently asked questions
Direct answers about securing LLM applications, agents, and the systems around them.
Agentic AI security is the practice of testing and hardening AI systems that can retrieve information, call tools, change data, or take actions. It covers the model and the runtime around it, including prompts, permissions, tools, memory, data, integrations, and deployment controls.
FailSafe tests prompt injection, indirect prompt injection, unsafe tool use, data leakage, RAG poisoning, tenant isolation, authentication and authorization, output handling, model access controls, and the security of connected APIs and infrastructure.
AI security assessments examine the model, agent, and integration-specific attack surface. SWARM adds continuous validation across the surrounding application, API, cloud, and infrastructure attack surface so fixes can be re-tested as the system changes.
Engagements can be mapped to the OWASP LLM Top 10, MITRE ATLAS, NIST AI RMF, SOC 2, ISO 27001, ISO 42001, EU AI Act evidence needs, and the controls that are relevant to the system's risk model.
Agentic testing should cover the full action loop: untrusted prompts and retrieved content, planning and memory, tool descriptions and arguments, agent identity, least-privilege permissions, multi-agent messages, recursive tool use, secrets, data exfiltration, supply-chain components, and human approval for high-impact actions.