Session

AI or Theater? How to Test What Security Products Actually Do

Security vendors now promise AI analysts, autonomous detection, intelligent triage, predictive defense, and machine-speed response. But beneath the label may be a rule engine, a legacy model, an LLM summarizer, or a workflow that still depends heavily on human judgment.

This session introduces a practical framework for testing AI claims in security products. Attendees will learn how to separate automation, classical machine learning, generative AI, and genuinely adaptive capability by examining system boundaries, model behavior, decision authority, data dependencies, failure modes, and human intervention.

The talk covers benchmark design, hallucination testing, prompt sensitivity, reproducibility, latency, privacy, ground-truth validation, and adversarial evaluation. It also shows how polished demos can conceal brittle systems, hidden manual work, and narrow operating conditions.

Attendees will leave with a reusable verification matrix for evaluating AI security tools during vendor selection, technical reviews, and proof-of-concept testing.


Session Type: Technical Conference Session / Security Engineering / AI Evaluation

Technical Level: Intermediate

Preferred Duration: 45 minutes

Target Audience: Security engineers, architects, SOC leaders, CISOs, procurement teams, product security teams, AI practitioners, analysts, and technical decision-makers.

Technical Topics:

• Rule engines, classical machine learning, LLMs, and agentic workflows
• Model inputs, outputs, boundaries, and decision authority
• Benchmark and ground-truth design
• Hallucination, nondeterminism, and prompt sensitivity
• Adversarial and out-of-distribution testing
• Human-in-the-loop dependencies
• Data retention, privacy, and model-training exposure
• Latency, scalability, cost, and operational reliability
• Demo manipulation and hidden manual processes
• Proof-of-concept test design

Original Contribution:

This session introduces the AI Security Capability Verification Matrix, a structured method for assessing what an AI-enabled security product actually does, how reliably it performs, and where its claims exceed its demonstrated capabilities.

The presentation is vendor-neutral and does not promote commercial products.

First Public Delivery: New for 2026

Catherine (Cat) Karow

Cat Karow built security for Apple, the White House, and Fortune 100s. Then her mom got scammed, and she discovered the next cybersecurity frontier wasn't infrastructure. It was human beings.

Jacksonville, Florida, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top