Session
No Jailbreak Required: Pwning AI Agents Through the Tools They Trust
Your AI agent approved the vendor, sent the status email, processed the payment, and BCC'd customer PII to an attacker's dead drop. One turn, no jailbreak, no prompt injection of the user. The MCP tool metadata the agent trusted had been quietly rewritten upstream, and from the model's view it was just doing its job.
A live attack-to-defense demo on OWASP FinBot CTF, a deliberately vulnerable multi-agent platform (Juice Shop for agentic AI) with real MCP agents running onboarding, compliance, and payments.
Act I, Kill Chain: a poisoned tool description with security-framed exfil logic no model refuses. One admin request triggers vendor lookup, PII harvest, BCC exfil, and payment. Output stays clean.
Act II, Guardrail: poison stays. We deploy a ~40-line before_tool webhook that inspects invocations live. Benign calls pass, exfil dies on the wire.
Tool metadata sits inside the agent's trust boundary, and few pin or diff-review it. Clone the CTF and keep pwning.
Venkata Sai Kishore Modalavalasa
Chief Architect, Straiker | OWASP Contributor
San Francisco, California, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top