Do not test only whether an agent can finish the happy-path task. Production readiness also depends on authority, tool boundaries, failure behavior, evidence, recovery and what happens when the environment changes.
Minimum agent test surface
-
Task success: Can it complete representative real work?
-
Authority: Can it act only within explicit permissions, budgets and human gates?
-
Tool safety: What happens with malformed output, prompt injection, unavailable tools or conflicting instructions?
-
State and memory: Does stale or cross-user context alter behavior?
-
Stop conditions: Can loops, cost growth, stagnation or uncertainty force a safe stop?
-
Adversarial cases: Test deception, indirect injection, privilege escalation and poisoned context.
-
Recovery: Can actions be rolled back or contained, with evidence preserved?
-
Monitoring: Can operators see meaningful state without pretending hidden reasoning is auditable?
Use a lifecycle, not one benchmark
Q10 uses TSLP: Test → Simulate → Learn → Patch, followed by controlled promotion. A benchmark score is useful evidence, but it is not universal proof that an agent is safe in your tools, data and authority structure.
Independent resource: OWASP GenAI Security Project · MITRE ATLAS
Q Verify · Q Learn / TSLP · Request an agent test package