Prompt injection happens when untrusted content influences an AI system to follow instructions the operator did not intend. It becomes more consequential when the model can call tools, access private data or take actions.
Direct vs indirect injection
Direct injection comes from the user or attacker interacting with the model. Indirect injection can be hidden inside a webpage, document, email, retrieved record or other content the AI processes.
Do not rely on one prompt as the security boundary
System prompts can guide behavior but they should not be the only protection around secrets, permissions or high-impact tools.
Layered controls
- Minimize privileges and tool scope.
- Separate trusted instructions from untrusted content.
- Validate tool inputs and outputs.
- Require approval for high-impact actions.
- Keep secrets outside model-visible context when possible.
- Use allowlists, budgets, rate limits and stop conditions.
- Log tool calls and preserve evidence.
- Test indirect-injection scenarios using real workflow content.
Independent resources
OWASP GenAI Security Project · MITRE ATLAS
Agent testing · Human control · Privacy & data · The Q Test