Session
Two Agents, Same Answer, One Is Wasting Your Money
Two agents answer the same question and both get it right. One made a single tool call. The other called a tool it did not need, then called the same one twice more. A pass or fail test scores both 100 percent, because it only sees the final answer, not the path. But those extra calls cost tokens, latency, and money every time it runs. This talk covers two ways to catch what pass or fail misses: score the answer with a clear rubric so results are consistent and explainable, and score the path by capturing every tool call and flagging duplicates, irrelevant tools, and wrong order. Together they catch the wasteful or unsafe runs that still land on a correct answer. You leave with working code for both.
Elizabeth Fuentes Leone
Developer Advocate
San Francisco, California, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top