Session
Route, Evaluate, Repeat: Cost-Aware LLM Architecture in Production
A production LLM system can return correct answers while wasting money, repeating failed actions, and hiding the change from its dashboard. One real agent spent six hours retrying the same work through an expensive model while every visible health indicator stayed green. Quality, cost, and behavior need to be observable together because any one of them can conceal trouble in the other two.
This session reconstructs that failure and builds the architecture the dashboard was missing. We expose why a model was chosen, when a retry stopped being reasonable, and what evidence justified an escalation. From there, we add routing and evaluation in layers, starting with cheap programmatic checks and bringing in expert judgment where it can change the decision. If you build LLM features that must survive production traffic and financial scrutiny, you'll leave knowing when routing earns its complexity and the minimum event trail every agent should produce before something goes wrong.
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top