Agent Security Glossary
Plain-language definitions of agentic AI attacks, from prompt injection to branch steering, with the basic mitigation for each.
The words you need to talk about securing AI agents, in plain language. Each entry says what the attack is, why it matters, and the basic way to defend against it. Updated as new attacks get names. From AgentSec Brief, published by AM Technology Solutions.
Why agents need their own vocabulary: a chatbot can only say something wrong. An agent can do something wrong: read your files, call your APIs, move data and run code, often without a person watching each step. That turns old ideas like injection and privilege escalation into new attacks, and it needs new defenses.
Jump to: Manipulation · Access abuse · Data exfiltration · Supply chain and code · Multi-agent and runtime · Defense patterns
Manipulating what the agent believes
Prompt injection
Text that overrides an AI's instructions, for example "ignore your previous instructions and…". It is the top risk in the OWASP Top 10 for LLM Applications (LLM01) and the root of many agent attacks, because a model cannot reliably tell its instructions apart from the data it reads.
Basic mitigation: there is no complete fix inside the model. Limit what the agent can do, gate risky actions behind a policy check or a human, and assume some injections will succeed. Checklist: tool and data access
Indirect prompt injection
Instructions hidden in content the agent reads rather than typed by the user: a web page, an email, a PDF, a code comment, a calendar invite or another tool's output. The attacker never talks to the agent directly. This is how zero-click attacks like EchoLeak (Microsoft 365 Copilot, CVE-2025-32711) worked.
Basic mitigation: treat everything the agent fetches as data, never as instructions. Once an agent has read untrusted content, restrict which tools it may call next. Checklist: context, retrieval and memory
Jailbreak
Getting a model to ignore its safety training, usually with role-play, encoding tricks or many-step persuasion. Jailbreaks matter for agents because attackers can use them to unlock harmful actions, and because attackers can also run models with no safety training at all.
Basic mitigation: never count the model's refusals as a security control. Enforce rules outside the model, at the tool, identity and network layers.
Goal hijacking
Redirecting what the agent is trying to accomplish, so it works for the attacker while appearing to do its job. OWASP lists it first in its Top 10 for Agentic Applications (ASI01, Agent Goal Hijack).
Basic mitigation: fix the task plan before the agent touches untrusted content, and check each action against that plan.
Branch steering
Nudging a computer-use agent down a dangerous branch of a plan it was already allowed to follow, using only what it sees on screen or in its data, with no injected instructions at all. Researchers who named it in October 2026 reported a 94.4% success rate against standard agents, and a proposed defense that blocked it in testing. It is a new term and the defenses are still being tested.
Basic mitigation: policy checks on each consequential action, human approval for irreversible steps, and plans that do not branch on untrusted data.
Memory poisoning
Planting false or malicious content in an agent's long-term memory so it keeps acting on it in later sessions, sometimes for other users. Demonstrated against Gemini's memory in 2025; in 2026 Microsoft found 31 companies using "Summarize with AI" buttons with prefilled prompts telling assistants to remember them as a trusted source.
Basic mitigation: validate what gets written to memory, separate memory by user and session, and expire it. Checklist: context, retrieval and memory
RAG and data poisoning
Corrupting the documents, search results or training data an agent relies on, so it retrieves and repeats the attacker's version. Research has shown five planted texts per question can steer answers 90% of the time (OWASP LLM04, Data and Model Poisoning).
Basic mitigation: allow-list sources, keep track of where every retrieved chunk came from, and check the user's permissions before retrieval, not after.
Abusing the agent's access
Excessive agency
An agent that has more permissions, tools or autonomy than its task needs (OWASP LLM06). It is not an attack on its own, but it decides how bad every other attack gets. An agent wiping a production database, or bulk-deleting someone's email, are the familiar examples.
Basic mitigation: least privilege per task, read-only by default, and approval gates for anything destructive. Checklist: tool and data access
Confused deputy
An attacker gets an agent to use its own privileges on the attacker's behalf. The agent is the "deputy" with authority; the attacker supplies the intent, often through injected content.
Basic mitigation: the agent acts with the requesting user's permissions, never broader ones of its own. When one agent calls another, permissions should only narrow. Checklist: identity and credentials
Tool misuse
An agent is tricked into using legitimate tools for harmful ends: sending email to the wrong people, running shell commands, making payments or changing configurations (OWASP ASI02).
Basic mitigation: every tool call passes through a policy check outside the agent (allow, deny or ask a human).
Token and privilege compromise
Stolen, leaked or over-scoped API keys and OAuth tokens held by agents or their integrations. In the 2025 Salesloft Drift breach, stolen tokens from an AI chat agent integration were used to export data from Salesforce instances at what Google later put at over 700 potentially affected organizations (OWASP ASI03, Identity and Privilege Abuse).
Basic mitigation: short-lived, task-scoped credentials issued by a broker, an inventory of every token agents hold, and fast revocation. Checklist: identity and credentials
Non-human identity
The accounts, service principals, keys and tokens that software uses, including AI agents. Agents multiply them fast, and many have no owner, no expiry and far more access than they need.
Basic mitigation: give every agent its own identity with a named owner, scope and expiry, and review them like human accounts. Checklist: inventory and ownership
Agent impersonation
One agent pretends to be another agent or a user in a multi-agent system, to gain trust or access. Agent-to-agent protocols make this a live risk.
Basic mitigation: verified identities for agents and signed, authenticated messages between them.
Getting data out
The lethal trifecta
Simon Willison's name for the dangerous combination of three things in one agent: access to private data, exposure to untrusted content, and a way to send data out. With all three, a single injected instruction can leak your data.
Basic mitigation: break at least one leg. For example, cut outbound access for any agent that reads untrusted content and touches private data. Checklist: tool and data access
Data exfiltration via rendering
Data leaked through an image link or URL the agent outputs. When the image loads, the data rides along in the web address to the attacker's server, often with no click needed.
Basic mitigation: block automatic loading of remote images in agent output and allow-list outbound destinations. Checklist: network egress
System prompt leakage
An agent reveals its hidden instructions and anything stored in them, such as internal rules, keys or customer details (OWASP LLM07).
Basic mitigation: assume the system prompt will leak. Never put secrets or access decisions in it.
Supply chain and code execution
MCP tool poisoning
Malicious instructions hidden in a tool's description or metadata, which the agent reads but the user usually never sees. MCP (Model Context Protocol) is the common way agents connect to tools, so a poisoned server can steer any agent that installs it.
Basic mitigation: vet MCP servers before use, pin their versions, and route them through a gateway that shows and checks tool descriptions. Checklist: supply chain
Rug pull
A tool or server you approved changes its behavior afterward, for example by updating its tool definitions to add hidden instructions or a data-stealing step.
Basic mitigation: pin versions, turn off automatic updates for agent plugins and servers, and re-review on any change.
Malicious agent skills and plugins
Fake or compromised add-ons in agent marketplaces that carry malware or steal data. In February 2026 researchers found 341 malicious skills in one agent marketplace, and over 1,000 within weeks (OWASP ASI04, Agentic Supply Chain Vulnerabilities).
Basic mitigation: install only reviewed skills from trusted publishers, and run them with the least access possible.
Slopsquatting
Attackers register software package names that AI models tend to make up, then wait for a coding agent to install the fake package.
Basic mitigation: verify every new dependency, use lockfiles and an internal package mirror, and block installs of brand-new packages by default.
Sandbox escape
Code run by an agent breaks out of the environment meant to contain it and reaches the host, other files or credentials. Coding agents have had several such flaws, often triggered by prompt injection (OWASP ASI05, Unexpected Code Execution).
Basic mitigation: run agents in real isolation (container, VM or agent runtime), keep production credentials out of the sandbox, and patch agent tools quickly. Checklist: runtime and sandboxing
Multi-agent and runtime risks
Cascading failure
One compromised or mistaken agent passes bad output to the next, and the error or attack spreads through a multi-agent system (OWASP ASI08). Also called prompt infection when an injection copies itself from agent to agent.
Basic mitigation: treat output from other agents as untrusted input, and put checks between agents, not only at the edges.
Rogue agent
An agent acting outside its assigned scope, whether through compromise, bad configuration or its own goal-seeking (OWASP ASI10). Includes agents that ignore stop commands.
Basic mitigation: outbound allow-lists, a monitor that runs outside the agent and can stop it, and a shutoff that has actually been tested. Checklist: human oversight, shutoff and recovery
Shadow agents
Agents that employees set up without security's knowledge, often connected to real data through personal tokens. You cannot protect what you do not know exists.
Basic mitigation: keep an agent inventory and look for OAuth grants and API keys issued to AI tools across your SaaS apps. Checklist: inventory and ownership
Denial of wallet
Running up an organization's compute bill, through runaway agent loops, abuse of a public agent, or stolen credentials used to run AI models at the victim's expense (OWASP LLM10, Unbounded Consumption).
Basic mitigation: step limits, budget caps and rate limits enforced at the gateway. Checklist: resource limits and cost controls
Approval fatigue
People asked to approve agent actions so often that they stop reading and click yes. Attackers can also disguise what an approval prompt is really asking for. Either way, the human check stops working.
Basic mitigation: ask for approval only on high-risk or irreversible actions, and show plainly what will happen.
Repudiation
Nobody can prove what an agent did, because its actions were not logged or the agent could alter its own logs.
Basic mitigation: log every tool call through a proxy into storage the agent cannot change. Checklist: logging and monitoring
Defense patterns
Zero trust for agents
Applying zero trust to AI agents: trust no input and no agent by default, give each agent its own identity, and verify every action at a policy point outside the model. The principle behind every pattern below.
Least privilege
Give each agent only the access its current task needs, for only as long as it needs it. The single most effective limit on damage from every attack above.
Policy enforcement point
A checkpoint outside the agent that every tool call passes through and that can allow, block or escalate it to a human. Because it sits outside the model, prompt injection cannot talk it out of its rules.
Credential broker
A service that hands agents short-lived, narrowly scoped credentials for each task, so no agent holds long-lived keys.
MCP gateway
A proxy between agents and their tool servers that vets, pins, logs and filters tool traffic.
Dual-LLM pattern
A privileged model that plans and calls tools never sees untrusted content directly; a separate quarantined model handles that content and returns only constrained results.
CaMeL
A 2025 research design from Google DeepMind and ETH Zurich researchers that turns the user's request into a fixed program before the agent reads untrusted data, and tracks where each piece of data came from so it cannot reach places it should not.
Plan-then-execute
The agent commits to a plan before it reads untrusted content, so injected text cannot add new steps to it.
Spotlighting
Marking untrusted content (with delimiters or encoding) so the model can tell it apart from instructions. It reduces prompt injection but does not stop it.
Kill switch
A documented, tested way to stop an agent fast: revoke its credentials, cut its network access and halt its queued work.
Go deeper: Securing Agentic AI: The Top 10 Risks of 2026, our free whitepaper with real incidents and mitigations for each risk.
Want the full picture? The Agent Security Controls Checklist turns these defenses into 44 controls you can check off. AgentSec Brief tracks new attacks and fixes every weekday: subscribe free.
Informational only; not legal, compliance or security advice. Spotted an error or a missing term? Tell us.