Session

AI Inference Best Practices: What Actually Works in Production

Running a model is easy. Running inference well is a lot harder.

What should you actually monitor? When should you batch requests? How do you choose an inference engine? Should you quantize? Scale up or scale out? How do you keep expensive GPUs busy without destroying latency? And what changes once you have more than one model, cluster, or type of accelerator?

This session walks through the practical best practices for running AI inference in production, from model and engine selection to GPU utilization, batching, caching, parallelism, routing, autoscaling, observability, and cost. We will look at the patterns that work, the common mistakes that look reasonable but hurt you later, and where the right answer depends on your workload.

With demos throughout, you will leave with a concrete checklist for designing, running, and improving a production inference stack.

Annie Talvasto

CNCF Ambassador & Sr. Manager at Upbound

New York City, New York, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top