Session

Trustworthy AI Agents: Test in CI, Enforce at Runtime

Companies are beginning to let AI agents search customer data, change settings, and call production tools. A secure model endpoint does not make those actions safe. An agent can choose the wrong tool, use a forbidden parameter, invent a value, or reach a correct answer through an unacceptable sequence of steps. This talk shows how to test those failures before release and how to stop them at runtime. We will use policy tests, grounded-answer checks, trajectory tests, and versioned tool contracts to build a practical control layer around an agent. Attendees will leave with a pattern they can adapt to their own CI pipeline and tool gateway.

What the audience should learn

- Why an answer-only evaluation misses unsafe agent behavior.
- How to turn tool policy, answer grounding, and expected trajectories into CI tests.
- Why schema validation alone does not catch tool drift, scope violations, or side-effect changes.
- How a runtime gate can return ALLOW, WARN, REVIEW, or BLOCK before a tool executes.
- How build-time tests and runtime controls complement each other.

Sachin Gupta

Technical Leader at eBay

San Jose, California, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top