Session

Break Your Agent Before Production Does

Your agent passes every test you wrote, because every test assumes the world behaves. Production does not: tools return wrong data, and some users push the agent on purpose. Two kinds of trouble, and you need both. Bad luck is chaos: a weather tool returns 12 degrees for Miami in June and the agent reports it as fact. A guardrail that range-checks the value the moment a tool returns catches it. Bad intent is red teaming: an attacker escalates over several turns until the agent leaks a stored card number. You generate these multi-turn attacks instead of scripting them, then score whether the agent held. Both are stochastic: run them many times and you get a correctness rate and a breach rate, not a single pass.


Outline: • The happy-path trap • Bad luck: chaos testing • Bad intent: red teaming • One run is not a measurement • Test both before you ship

Elizabeth Fuentes Leone

Developer Advocate

San Francisco, California, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top