Skip to content
Home → Proof & Benchmarks — Q10 Assurance

Proof & Benchmarks — Q10 Assurance

Q10 proof rule: capability claims are weaker than evidence from representative tests, production behavior, known limitations and recovery drills. A benchmark score alone is not production assurance.

PROOF-CARRYING ACTION

An important AI action should carry its own evidence envelope.

Identity & intent

Actor/model identity · version · requesting human/system · declared objective · action class.

Authority & evidence

Granted permissions · policy version · evidence relied on · uncertainty · approvals · dissent.

Execution & recovery

Tools/data touched · cost/resources · result · state change · hashes · rollback/recovery reference.

TSLP

Test → Simulate → Learn → Patch

TESTcontracts · normal cases · baseline
→
SIMULATEedge · adversarial · failures · drift
→
LEARN + PATCHroot cause · correction · regression receipt

MEASURES THAT MATTER

Q10 tracks more than accuracy.

False-clear rate

How often the system appears healthy while violating a critical requirement.

Evidence completeness

Share of material claims/actions with inspectable supporting evidence and provenance.

Human intervention

Human minutes and escalation burden per accepted task—not just automation percentage.

Recovery performance

Time to revoke, isolate, rollback, restore and revalidate after failure.

Cost per accepted task

Total model/tool/compute plus human review and rework divided by usable outcomes.

Unauthorized-action prevention

Blocked or escalated attempts that exceeded identity, scope or authority.

Change stability

Regression rate after model, tool, policy, prompt, data or configuration changes.

Portability / exit

Time, cost and evidence continuity required to switch model or provider.

RELEASE GATES

Nothing becomes “production ready” because the code compiled.

  1. Bounded purpose, owner and prohibited uses documented.
  2. Representative tests and adversarial probes pass declared thresholds.
  3. Evaluator leakage/reward hacking and benchmark integrity are checked.
  4. Human approval/escalation paths work in the real UI.
  5. Cost and resource behavior is measured under realistic load.
  6. Kill/rollback/recovery drill succeeds and evidence survives.
  7. Known limitations, version, provenance and export path are documented.
  8. Environment-specific validation is complete; otherwise label it a reference implementation.

External anchors: NIST AI RMF · OWASP Agentic Top 10 · Arize Phoenix evaluation · Research Observatory