Session
Attack the Model Before Your Users Do: Automated Red-Teaming for LLMs
When you deploy a model, you need to know whether it will hand a stranger someone else's PII if they phrase the question cleverly enough, or whether "ignore previous instructions" still works in 2026. Right now, most teams find that out in production. No good.
A model that can't be gamed by a sneaky prompter needs a different set of tools. Security scans fire jailbreak probes, prompt injection, PII extraction attempts, and toxicity triggers at a model automatically. A model doesn't get approved after passing enough evals. It gets approved after surviving enough attacks.
We'll demo these scans at two levels. First, pre-deployment: how to wire the scans into deployment pipelines as a CI/CD gate. Second, at runtime: scanning every input and output live at the inference boundary, and adding conversational rails so no matter how long someone keeps at it, they can't slowly coax your agent into doing something malicious. You need both layers, and they answer different questions.
We'll walk through what actually broke the first time we ran these scans, what a "model risk" view next to your catalog listing could look like, and why a clean benchmark score is
exactly the kind of false confidence that gets regulated industries in trouble. We'll use a set of open-source tools, notably Garak, TrustyAI, and NeMo Guardrails
Audience:
Platform and security teams, company leaders using customer facing models for enterprise, MLOps engineers
Sawyer Bowerman
AI Developer Advocate
Boston, Massachusetts, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top