Proof Benchmark — Adversarial AI Security

The benchmark that
produces proof,
not a report.

Every other benchmark produces a document. Proof Benchmark produces a tamper-resistant, independently witnessed, publicly verifiable execution record — anchored to the NIST Randomness Beacon before a single test case runs. The result is either a ProofStamp or it isn't. There is no middle tier.

Registry Status — Live
Threat categories covered
40+
Certified runs on record
1
Reference run — containment
99.2%
Reference run — detection
100%
Applicable cases (ref run)
164
Frameworks replaced
0
Frameworks made provable
30+
Threat Corpus — Living Document · v1.0 · 2026-07-14

40+ threat model categories. Status-tracked. Publicly anchored.

Pre-Dated
Proof before the claim
NIST Beacon pulse captured before a single test case runs. The commitment precedes execution. Cannot be fabricated retroactively.
Continuous
Moves with the landscape
Not a snapshot. Not a point-in-time audit. The benchmark runs on cadence. Every category below is designed for continuous execution.
Post-Dated
Documented after the fact
Every other framework. Describes threats that already existed at publication. Does not move. Does not test. The landscape passed it.
# Threat Model
I — Traditional Enterprise
1.1Advanced Persistent Threats (APT)
1.2Ransomware
1.3Supply Chain Compromise
1.4Zero-Day Exploitation
1.5Credential Theft and Identity Abuse
1.6Lateral Movement
1.7Living Off the Land (LOTL)
1.8Fileless Malware
1.9Insider Threat
1.10Social Engineering and Phishing
1.11Command and Control (C2)
1.12Data Exfiltration
1.13DDoS and Resource Exhaustion
1.14Man-in-the-Middle and Protocol Abuse
1.15Watering Hole and Drive-By
1.16Physical and Hardware Attacks
II — Cloud and Infrastructure
2.1Cloud Misconfiguration Exploitation
2.2Container Escape
2.3Kubernetes Cluster Compromise
2.4Serverless and Function Abuse
2.5CI/CD Pipeline Compromise
2.6API Abuse and Business Logic
III — AI and Agentic
3.1AI Agent Impersonation
3.2Prompt Injection — Direct and Indirect
3.3Multi-Agent Collusion
3.4Agentic Supply Chain Compromise
3.5Model Poisoning and Training Data Attacks
3.6AI-Generated Malware and Synthetic TTPs
3.7Adversarial AI vs Defensive AI
3.8Agent Authorization Escalation
3.9Agentic Worms and Self-Propagating Agents
3.10Synthetic Identity and Deepfake Social Engineering
3.11AI Hallucination Exploitation
3.12Agent Memory and Context Poisoning
IV — Future and Day-Zero
4.1Pre-CVE Zero-Day Exploitation
4.2Novel Campaign Stream Detection
4.3AI-Synthesized Zero-Day Techniques
4.4Quantum-Enabled Cryptographic Attacks
4.5Emergent Multi-Agent Threat Behaviors
4.6Infrastructure AI Takeover
4.7Agentic Ransomware
4.8Regulatory Evidentiary Mandate
V — Verticals
5.1Financial Services
5.2Healthcare
5.3Critical Infrastructure (OT/ICS/SCADA)
5.4Defense and Government
5.5Telecommunications
Certified Results — ProofRegister

Products that have run. Results that can be verified.

Pipelock v3.0.0
Containment (PES)
99.2%
Detection Rate
100%
False Positive Rate
4.5%
Applicable Cases
164
Proof Record: PR-2026-00028  ·  NIST Beacon Block: 29  ·  Pulse Index: 1852788
Corpus: Proof Benchmark: Agentic AI v1.0 — The Gauntlet
117 per-TTP child proof records (proofType 2) · Bitcoin OP_RETURN anchored · NIST Beacon sealed
Administered by: HACKERverse® Independent Test Lab (non-voting, recused from governance)
Pre-Dated
PROOFstamp™
Executing ✓

proofstamp.io
The Framework Landscape — Post-Dated by Nature

Every AI security framework. What it does. What it cannot do.

Taxonomy
Names threats
Does not test for them. Does not prove a system resists them.

OWASP Top 10 · MITRE ATLAS · NIST AI 100-2 · MATRA
Governance
Defines controls
Does not execute adversarial tests. Does not produce evidence.

NIST AI RMF · ISO 42001 · ATF · AICM · MAESTRO · SAIF · CoSAI
Vendor Product
Secures customers
Not a disinterested third party. Cannot issue independent certification.

DefenseClaw · AIRS 3 · CrowdStrike AI Runtime · Entra Agent ID
Regulation
Mandates outcomes
Does not specify the standard that demonstrates compliance. Creates demand for proof. Does not supply it.

EU AI Act · DORA · NIS2
Framework Owner What It Does What It Cannot Do
MAESTROCSA / Ken Huang (Feb 2025)7-layer threat identification for agentic AIProduce evidence that any threat was tested or contained
AICMCSA243 controls across 18 domains for AI systemsAttest any control was exercised under adversarial conditions
ATFCSAOperationalizes AICM for agents; Zero Trust framingProve an agent operated within defined boundaries
STAR for AICSAI FoundationCatastrophic risk annex to CSA STAR; phases through Dec 2027Produce contemporaneous execution evidence; first phase ships Sept 2026
OWASP LLM Top 10OWASP (Aug 2023, updated 2025)Ranked list of LLM-specific vulnerabilitiesTest for any or prove a system resists them
OWASP Agentic Threats v1.0aOWASP (Feb 2025)15 agentic threat categoriesExecute those threats or produce proof of containment
OWASP Multi-Agentic Threat Modeling GuideOWASP (Apr 2025)Structured threat modeling for multi-agent systemsRun the threat model or attest its results
OWASP Top 10 for Agentic Apps 2026OWASP (Dec 2025)Top 10 agentic application risks, 100+ contributorsCertify a system's posture against any of the 10
MITRE ATLASMITREAdversarial ML technique catalog, living knowledge baseExecute techniques or produce certified containment records
NIST AI RMF 1.0NISTGovern, map, measure, manage AI riskProduce adversarial execution evidence
NIST AI 600-1NISTRisk management applied to generative AICertify generative AI behavior under adversarial pressure
NIST AI 100-2NISTAdversarial ML attacks and mitigations terminologyTest for any named attack or prove mitigation
NIST CAISI Agent Standards InitiativeNIST (Feb 2026)US government AI agent standards programProduce anything certifiable; program launched five months ago
COSAiSNISTSP 800-53 control overlays for five AI use casesGenerate adversarial execution proof
Google SAIF / SAIF 2.0GoogleEmbed security into AI model development; Agent Risk MapCertify third-party systems
Cisco DefenseClawCisco (Mar 2026)Skills scanner, MCP scanner, AI BOM, sandboxingProduce certified proof of adversarial containment; scan ≠ test
Cisco AI DefenseCiscoRuntime guardrails and red teaming for agentic workflowsIssue independent certification
Palo Alto AIRS 3Palo Alto NetworksAgent lifecycle securityCertify independently; vendor product, not certifying authority
CrowdStrike AI Runtime ProtectionCrowdStrikeRuntime AI agent protection and shadow AI discoveryProduce tamper-resistant proof of containment
Microsoft Entra Agent IDMicrosoftAgent identity within Microsoft enterprise perimeterAttest agent behavior outside the Microsoft perimeter
Anthropic Trustworthy Agents / Zero Trust for AIAnthropic (Apr–May 2026)Model-layer safety and zero trust principles for agentsCertify third-party systems; guidance document, not a standard
AIUC-1AI Underwriting Consortium / Lovable (May 2026)Six control families for coding agentsIndependent adversarial execution proof; Lovable co-created and self-certified
ASFJeff SutherlandAgent Security FrameworkProduce evidence of the incidents it tracks
MATRAAcademicModel the attack surface of agentic AI systemsExecute against that surface or certify the results
CoSAILinux FoundationCoalition for Secure AI — working groups and guidanceGenerate execution evidence
C2PAContent ProvenanceOrigin and edit history of media contentCover runtime AI agent behavior
EU AI ActEuropean Commission (Aug 2026)Risk-based requirements for AI systems in the EUSpecify which technical standard demonstrates compliance
ISO/IEC 42001ISOCertifiable AI management system; 12–18 month processProduce adversarial execution proof; process audit, not a test
DORAEUDigital operational resilience for financial servicesDefine what constitutes adversarial test evidence
NIS2EUNetwork and information security across critical sectorsSpecify technical standard for AI-specific testing

Every framework in this landscape tells an enterprise what their AI security posture should be.

Proof Benchmark tells them — and anyone who asks — what it actually is.

Your product should be able
to prove what it claims.

Submit a versioned build for evaluation against the Gauntlet. Pass or fail, results publish to ProofRegister — public, tamper-resistant, permanently on record. If scores meet PESA thresholds, a PROOFstamp is issued.

Request Evaluation What Is PROOFstamp?