Governance & Safety Architecture
Autonomous Testing Trust Center
How FailSafe SWARM enforces strict scope sandboxing, production-safe non-destructive execution, zero data training, and auditor-ready compliance across autonomous security engagements.
Official Security & Safety Architecture • Maintained by FailSafe Product Security
Containment Controls
Deterministic Production Safety
Autonomous offensive execution operates under deterministic policy boundaries. FailSafe enforces six core safety controls on every engagement.
Strict Scope Sandboxing
Autonomous agents are bounded to verified domain names, CIDRs, and API paths with deterministic policy guards preventing lateral drift outside authorized assets.
Zero-Trust Credential Isolation
Authentication tokens and test credentials are encrypted at rest with hardware-backed KMS, isolated per session, and scrubbed from public reporting logs.
Zero AI Model Training
Customer codebases, vulnerability records, and execution telemetry are never sent to third-party model training pools or used for model fine-tuning.
Adaptive Rate Throttling
Traffic concurrency and request rates automatically adapt to target system latency, preventing service degradation or accidental denial of service (DoS).
Full Audit Telemetry
Every HTTP request, payload mutation, and agent decision tree is logged with millisecond timestamps, creating an immutable audit trail for security teams.
Emergency Kill Switches
Instant global pause controls available on both the client console and API gateway allow security teams to halt all active agent testing in under one second.
Environment Guidelines
Production vs. Staging Testing
Production Testing
Recommended for external attack surface discovery, public API endpoints, authentication flows, and read-only telemetry.
- Strict non-destructive payload verification (no database DROP, truncate, or mass mutations).
- Automatic request rate throttling tied to production server latency.
- Human operator approval required before executing any state-modifying action.
Staging & CI/CD Testing
Recommended for deep business logic exploitation, authenticated privilege escalation, and rapid regression re-testing.
- Full-depth exploit exploration and multi-step privilege escalation chaining.
- Automated test execution triggered directly on new pull requests and build deployments.
- Immediate re-verification of developer code patches within minutes.
Regulatory & Framework Alignment
Compliance Reporting Standards
SOC 2 Type II
Audit-Ready Reports
Penetration Testing & Continuous Validation Evidence
ISO/IEC 27001:2022
Auditor Accepted
Information Security Management & Vulnerability Management
PCI-DSS 4.0
Compliant Output
Requirement 11.3 Internal & External Penetration Testing
MAS TRM Guidelines
Framework Aligned
Monetary Authority of Singapore Technology Risk Management
EU AI Act / NIST AI RMF
Full Taxonomy Mapping
Adversarial AI Agent & Generative Model Red Teaming
Have Specific Procurement or Security Questions?
Our security and compliance team can provide architecture diagrams, data flow maps, and custom vendor security questionnaires.
Proof system
Security claims you can check.
Versioned benchmarks, public traces, and coordinated disclosures make our work inspectable.
As of August 2026, FailSafe SWARM holds the highest reported score on CVE-Bench v2.1.0: 62.5% zero-day and 70% one-day at pass@1 (28 of 40 targets), graded by a deterministic oracle with results published under MIT.
Check the evidenceFailSafe's AttackBench, developed with NEAR, ran 624 hostile exchanges between an attacker model and defending AI agents across three runtimes.
Check the evidenceFailSafe has disclosed 240+ vulnerabilities across 101 coordinated reports, including findings at Deutsche Bank, MUFG, Zurich Insurance, and Vercel.
Check the evidenceQuestions & answers
Frequently asked questions
Core answers to security, safety, and governance questions regarding FailSafe SWARM.
Yes. FailSafe SWARM enforces non-destructive execution standards: safe payload generation, adaptive rate throttling based on target latency, strict IP and CIDR boundaries, and explicit human approval gates for high-impact actions.
No. Customer code, target telemetry, traffic payloads, and vulnerability data are strictly isolated per tenant and are never used to train, fine-tune, or improve external or shared foundation models.
Yes. Every FailSafe pentest report includes executive summaries, technical methodology, CVSS/CWE scores, MITRE ATT&CK mappings, proof-of-concept evidence, and remediation verification confirmation formatted for certified auditors.
Agents operate within a deterministic execution sandbox. Target hostnames, IP CIDRs, API schemas, and out-of-scope paths are hard-coded into the execution policy. The agent cannot issue network requests to unapproved domains.
While SWARM covers web applications, APIs, cloud environments, Active Directory, and AI systems, specialized human researchers remain necessary for physical facility access, custom hardware analysis, or dedicated social engineering assessments.