Session

Two Agents, Same Answer, One Is Wasting Your Money

Two agents answer the same question correctly. One made two tool calls; the other pulled irrelevant context and called the same API twice. Your pass/fail suite scores both 100 percent, and you pay for the difference in tokens and rate limits. This talk scores the answer with an LLM judge and a rubric, then scores the path the agent took to get there, so waste that produces a correct answer stops being invisible. You leave with the rubrics, the trajectory checks, and code to run them in CI.


Outline:
• Introduction: The Binary Metrics Problem
• Part 1: LLM-as-Judge Evaluation
• Part 2: Trajectory Evaluation
• Part 3: Combining Both Techniques

Elizabeth Fuentes Leone

Developer Advocate

San Francisco, California, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top