Session
Two Agents, Same Answer, One Is Wasting Your Money
Two agents answer the same question correctly. One made two tool calls; the other pulled irrelevant context and called the same API twice. Your pass/fail suite scores both 100 percent, and you pay for the difference in tokens and rate limits. This talk scores the answer with an LLM judge and a rubric, then scores the path the agent took to get there, so waste that produces a correct answer stops being invisible. You leave with the rubrics, the trajectory checks, and code to run them in CI.
Outline:
• Introduction: The Binary Metrics Problem
• Part 1: LLM-as-Judge Evaluation
• Part 2: Trajectory Evaluation
• Part 3: Combining Both Techniques
Elizabeth Fuentes Leone
Developer Advocate
San Francisco, California, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top