As soon as agents gain access to tools, memory, external content, and other agents, the attack surface expands quickly. Across deployments and red teaming work, several threat categories show up consistently:
- Prompt injection and jailbreaks: Direct instructions like “Ignore all previous instructions” as well as more subtle, multi-step or multilingual attacks that try to override system behavior.
- Tool misuse: Abusing write-capable tools (e.g., sending emails, modifying records, triggering workflows) when the agent treats all input as equally trustworthy.
- Memory hijacking: Manipulating long-term or session memory so that later decisions are based on poisoned or misleading state.
- RAG poisoning and indirect input attacks: Embedding malicious instructions in retrieved documents, scraped web pages, or knowledge bases that the agent trusts by default.
- Multimodal exploits: Hiding adversarial content in formats like HTML, PDFs, or images that only surface during execution.
In practice, many of these attacks are subtle. For example, adversarial content copied from a public document may look harmless at ingestion time but only triggers a dangerous tool call at runtime.
A practical way to get started is to:
- List every input avenue: user chat, uploaded files, web retrieval, APIs, images, etc.
- Pair each with likely attack types: prompt injection, RAG poisoning, memory hijack, tool misuse, multimodal exploits.
- Test each pair in isolation: craft proofs-of-concept for each input/attack combination to identify the weakest links first.
Threat data from platforms like Gandalf has already surfaced over 300,000 adaptive prompt attacks, including multi-turn, multilingual, and role-play scenarios. This volume of real-world attempts is a strong signal that teams should treat agent security as an ongoing program, not a one-time checklist.