AI reliability is not only about preventing failure. It is also about limiting damage and recovering when prevention fails.
A practical recovery sequence
-
Stop or contain. Prevent further unsafe or unwanted actions.
-
Preserve evidence. Keep relevant logs, prompts, tool calls, versions, approvals and environmental state.
-
Assess scope. Determine what changed, who or what was affected and whether downstream systems propagated the issue.
-
Roll back or restore. Revert to a known state where technically and operationally possible.
-
Correct. Patch the model interaction, data, policy, tool, integration or workflow that contributed to failure.
-
Retest. Reproduce the incident and test nearby failure modes.
-
Promote deliberately. Restore authority only after the required evidence and human gate are satisfied.
Plan for provider failure too
Recovery should cover outages, account lockout, model deprecation, pricing or terms changes, corrupted memory, deleted data and loss of a critical operator—not only bad model answers.
Q Recovery & Continuity · Human control · Evidence & status