Session

The Output Looks Right. Where’s the Evidence?

LLMs are exceptionally good at producing outputs that look complete, coherent, and convincing. That creates a dangerous shortcut: we read the result, it makes sense, and we continue working as though it were correct.

Sometimes that is enough. Sometimes it is not.

For higher-impact work, plausibility is not evidence. We may need to know that the model used the right information, performed the necessary checks, considered relevant constraints, and produced an outcome that satisfies the actual requirement.

The challenge is not to verify every token an LLM generates. It is to decide what evidence this particular task requires before accepting the result.

In this session, I’ll share a practical framework we use to define evidence requirements for AI-generated work. We’ll examine when human review is sufficient, when stronger evidence is necessary, and how to make evidence part of the workflow instead of an investigation performed only after something goes wrong.


Target audience: AI engineers, software engineers, product teams, tech leads, and anyone responsible for reviewing AI-generated work.
Level: Intermediate.
Preferred duration: 30–45 minutes.
Format: Practical framework with real examples from AI workflows.
Prerequisites: Familiarity with using LLMs for consequential development, research, or product tasks.
Source: Lessons from designing validation and evidence requirements around production LLM systems.

Haberman Michael

3× Founder & CTO | Building Reliable Software in the AI Era

Tel Aviv, Israel

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top