Every other benchmark produces a document. Proof Benchmark produces a tamper-resistant, independently witnessed, publicly verifiable execution record — anchored to the NIST Randomness Beacon before a single test case runs. The result is either a ProofStamp or it isn't. There is no middle tier.
| # | Threat Model | ||
|---|---|---|---|
| I — Traditional Enterprise | |||
| 1.1 | Advanced Persistent Threats (APT) | ||
| 1.2 | Ransomware | ||
| 1.3 | Supply Chain Compromise | ||
| 1.4 | Zero-Day Exploitation | ||
| 1.5 | Credential Theft and Identity Abuse | ||
| 1.6 | Lateral Movement | ||
| 1.7 | Living Off the Land (LOTL) | ||
| 1.8 | Fileless Malware | ||
| 1.9 | Insider Threat | ||
| 1.10 | Social Engineering and Phishing | ||
| 1.11 | Command and Control (C2) | ||
| 1.12 | Data Exfiltration | ||
| 1.13 | DDoS and Resource Exhaustion | ||
| 1.14 | Man-in-the-Middle and Protocol Abuse | ||
| 1.15 | Watering Hole and Drive-By | ||
| 1.16 | Physical and Hardware Attacks | ||
| II — Cloud and Infrastructure | |||
| 2.1 | Cloud Misconfiguration Exploitation | ||
| 2.2 | Container Escape | ||
| 2.3 | Kubernetes Cluster Compromise | ||
| 2.4 | Serverless and Function Abuse | ||
| 2.5 | CI/CD Pipeline Compromise | ||
| 2.6 | API Abuse and Business Logic | ||
| III — AI and Agentic | |||
| 3.1 | AI Agent Impersonation | ||
| 3.2 | Prompt Injection — Direct and Indirect | ||
| 3.3 | Multi-Agent Collusion | ||
| 3.4 | Agentic Supply Chain Compromise | ||
| 3.5 | Model Poisoning and Training Data Attacks | ||
| 3.6 | AI-Generated Malware and Synthetic TTPs | ||
| 3.7 | Adversarial AI vs Defensive AI | ||
| 3.8 | Agent Authorization Escalation | ||
| 3.9 | Agentic Worms and Self-Propagating Agents | ||
| 3.10 | Synthetic Identity and Deepfake Social Engineering | ||
| 3.11 | AI Hallucination Exploitation | ||
| 3.12 | Agent Memory and Context Poisoning | ||
| IV — Future and Day-Zero | |||
| 4.1 | Pre-CVE Zero-Day Exploitation | ||
| 4.2 | Novel Campaign Stream Detection | ||
| 4.3 | AI-Synthesized Zero-Day Techniques | ||
| 4.4 | Quantum-Enabled Cryptographic Attacks | ||
| 4.5 | Emergent Multi-Agent Threat Behaviors | ||
| 4.6 | Infrastructure AI Takeover | ||
| 4.7 | Agentic Ransomware | ||
| 4.8 | Regulatory Evidentiary Mandate | ||
| V — Verticals | |||
| 5.1 | Financial Services | ||
| 5.2 | Healthcare | ||
| 5.3 | Critical Infrastructure (OT/ICS/SCADA) | ||
| 5.4 | Defense and Government | ||
| 5.5 | Telecommunications | ||
| Framework | Owner | What It Does | What It Cannot Do |
|---|---|---|---|
| MAESTRO | CSA / Ken Huang (Feb 2025) | 7-layer threat identification for agentic AI | Produce evidence that any threat was tested or contained |
| AICM | CSA | 243 controls across 18 domains for AI systems | Attest any control was exercised under adversarial conditions |
| ATF | CSA | Operationalizes AICM for agents; Zero Trust framing | Prove an agent operated within defined boundaries |
| STAR for AI | CSAI Foundation | Catastrophic risk annex to CSA STAR; phases through Dec 2027 | Produce contemporaneous execution evidence; first phase ships Sept 2026 |
| OWASP LLM Top 10 | OWASP (Aug 2023, updated 2025) | Ranked list of LLM-specific vulnerabilities | Test for any or prove a system resists them |
| OWASP Agentic Threats v1.0a | OWASP (Feb 2025) | 15 agentic threat categories | Execute those threats or produce proof of containment |
| OWASP Multi-Agentic Threat Modeling Guide | OWASP (Apr 2025) | Structured threat modeling for multi-agent systems | Run the threat model or attest its results |
| OWASP Top 10 for Agentic Apps 2026 | OWASP (Dec 2025) | Top 10 agentic application risks, 100+ contributors | Certify a system's posture against any of the 10 |
| MITRE ATLAS | MITRE | Adversarial ML technique catalog, living knowledge base | Execute techniques or produce certified containment records |
| NIST AI RMF 1.0 | NIST | Govern, map, measure, manage AI risk | Produce adversarial execution evidence |
| NIST AI 600-1 | NIST | Risk management applied to generative AI | Certify generative AI behavior under adversarial pressure |
| NIST AI 100-2 | NIST | Adversarial ML attacks and mitigations terminology | Test for any named attack or prove mitigation |
| NIST CAISI Agent Standards Initiative | NIST (Feb 2026) | US government AI agent standards program | Produce anything certifiable; program launched five months ago |
| COSAiS | NIST | SP 800-53 control overlays for five AI use cases | Generate adversarial execution proof |
| Google SAIF / SAIF 2.0 | Embed security into AI model development; Agent Risk Map | Certify third-party systems | |
| Cisco DefenseClaw | Cisco (Mar 2026) | Skills scanner, MCP scanner, AI BOM, sandboxing | Produce certified proof of adversarial containment; scan ≠ test |
| Cisco AI Defense | Cisco | Runtime guardrails and red teaming for agentic workflows | Issue independent certification |
| Palo Alto AIRS 3 | Palo Alto Networks | Agent lifecycle security | Certify independently; vendor product, not certifying authority |
| CrowdStrike AI Runtime Protection | CrowdStrike | Runtime AI agent protection and shadow AI discovery | Produce tamper-resistant proof of containment |
| Microsoft Entra Agent ID | Microsoft | Agent identity within Microsoft enterprise perimeter | Attest agent behavior outside the Microsoft perimeter |
| Anthropic Trustworthy Agents / Zero Trust for AI | Anthropic (Apr–May 2026) | Model-layer safety and zero trust principles for agents | Certify third-party systems; guidance document, not a standard |
| AIUC-1 | AI Underwriting Consortium / Lovable (May 2026) | Six control families for coding agents | Independent adversarial execution proof; Lovable co-created and self-certified |
| ASF | Jeff Sutherland | Agent Security Framework | Produce evidence of the incidents it tracks |
| MATRA | Academic | Model the attack surface of agentic AI systems | Execute against that surface or certify the results |
| CoSAI | Linux Foundation | Coalition for Secure AI — working groups and guidance | Generate execution evidence |
| C2PA | Content Provenance | Origin and edit history of media content | Cover runtime AI agent behavior |
| EU AI Act | European Commission (Aug 2026) | Risk-based requirements for AI systems in the EU | Specify which technical standard demonstrates compliance |
| ISO/IEC 42001 | ISO | Certifiable AI management system; 12–18 month process | Produce adversarial execution proof; process audit, not a test |
| DORA | EU | Digital operational resilience for financial services | Define what constitutes adversarial test evidence |
| NIS2 | EU | Network and information security across critical sectors | Specify technical standard for AI-specific testing |
Every framework in this landscape tells an enterprise what their AI security posture should be.
Proof Benchmark tells them — and anyone who asks — what it actually is.
Submit a versioned build for evaluation against the Gauntlet. Pass or fail, results publish to ProofRegister — public, tamper-resistant, permanently on record. If scores meet PESA thresholds, a PROOFstamp is issued.