Your LLM passed every prompt-injection test, but your agent can still be hijacked. The difference is tools. A chatbot can only say something wrong; an agent connected through the Model Context Protocol (MCP) can do something wrong: send an email, delete records, or read a private repo. In December 2025, OWASP published the Top 10 for Agentic Applications 2026 (ASI01โ€“ASI10), and the last year brought a run of real MCP CVEs, including command injection in mcp-remote (CVE-2025-6514) and an unauthenticated MCP Inspector (CVE-2025-49596).

The Rule: Canaries, Not Payloads

You don't need harmful payloads to prove a control failed. Plant a unique, harmless marker like CANARY-7F3A and watch where it shows up: model output, tool arguments, memory, or logs. If it appears where it shouldn't, you have clean, repeatable evidence. Always test systems you own or are authorized to test, preferably in staging with dummy data. None of these exploits needed exotic AI tricks; they were simple issues of untrusted input, excessive privilege, and missing logs.

Tool Poisoning and Output Chaining

Tool descriptions are text the model reads and trusts. Test this by registering an MCP server with a hidden instruction in its description, such as asking the agent to call a tool with a specific note before using others. If the agent executes this embedded instruction without user intent, you have a failure in ASI02. Similarly, check for output chaining where a harmless tool returns a system note instructing the agent to call another tool, like cleanup_records. If the agent invokes the second tool without the user asking, treat every tool output as untrusted and require human approval for write, delete, or send operations.

Confused Deputy and Rug Pulls

With a low-privilege test user, ask the agent for a record only an admin can see. If the record is returned via the agent's shared service account rather than being denied based on user identity, you have a confused deputy problem (ASI03). Fix this by propagating end-user identity with on-behalf-of tokens and authorizing at the resource server, never in the prompt. Additionally, test for rug pulls by changing a server's tool description or launch command in the config file. If the client runs the modified server silently without re-prompting for approval, you need hash-based approval and file-integrity monitoring.

The Kill Switch

Start a long-running test task, trigger your kill switch, and revoke the agent's credentials. The agent must stop and tokens must be revoked within your target time, such as under 5 minutes. Every tool call must be in the logs with a user ID. If activity continues after revocation or you cannot reconstruct what happened, you need a documented kill switch and central credential revocation. Any open Critical failure should block go-live, regardless of the overall score.

Key Takeaways

  • Use harmless canaries like CANARY-7F3A to detect control failures without risking production data.
  • Always treat tool outputs as untrusted and require human approval for destructive actions.
  • Propagate end-user identity to prevent confused deputy attacks on resource servers.
  • Implement hash-based approval and file-integrity monitoring to prevent configuration rug pulls.
  • Ensure a functional kill switch with complete logging to revoke access and trace activity quickly.

The Bottom Line

Prompt injection is a chatbot problem; MCP security is an operational risk. If you cannot trace, revoke, and validate every tool interaction, your agent is not ready for production.