Session

Break me if you can: Evaluating safety robustness in AI models

AI safety is often framed as an ethical question, but for businesses shipping real products, it is also a matter of security, trust, and legal integrity. Unsafe model behavior can expose secrets, amplify bias, enable abuse, and create serious reputational risk. This talk gives developers and product teams a practical introduction to AI alignment and AI safety through the lens of how systems actually fail in production.

From there, we examine how the AI red teaming field is moving from manual probing to automated adversarial evaluation at scale, including static test sets, agentic jailbreaking, and optimizer-driven attack discovery. They close with a practical threat model, concrete mitigations, and an overview of the current AI security landscape, giving developers a grounded framework for building safer, more secure AI systems.

Áron Erdélyi

Senior Consultant @ TNG | Full-Stack Developer | Generative AI Specialist

Austin, Texas, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top