Researcher Discloses MCP Agent-to-Agent Injection Attacks
Proof-of-concept exploits abuse trust between internal AI agents to spread malicious instructions, according to Ars Technica.
An independent researcher has disclosed proof-of-concept attacks that exploit trust gaps in the Model Context Protocol (MCP) to spread malicious instructions between internal AI agents at several organizations. The researcher, Syed Anas Mohiuddin, detailed the technique, which is a form of prompt injection that targets a specific agent rather than the large language model powering it.
According to Ars Technica, the attacks abuse trust relationships among agents, allowing harmful instructions to propagate from one agent to others within a targeted network. In many cases, crafted prompts led to server-side request forgery, the researcher found.
Mohiuddin tested the technique against agents from Google, JP Morgan Chase, Weviate, Rapid7, the French government's interministerial digital directorate, and the US federal government.
Earlier disclosures
The new proof-of-concept follows prior acknowledgments of similar risks. In the past five months, Google and four other organizations have acknowledged vulnerabilities that exploit one agent inside a targeted network to spread harmful instructions to other internal agents.
Quick answers
What is the Model Context Protocol?
MCP is a protocol used to connect AI agents and tools. The disclosed attacks exploit trust gaps in how agents communicate through it.
Who disclosed the attacks?
Independent researcher Syed Anas Mohiuddin demonstrated the proof-of-concept attacks.
Which organizations' agents were tested?
Agents from Google, JP Morgan Chase, Weviate, Rapid7, the French government's interministerial digital directorate, and the US federal government were tested.