Session

From Platform Engineer to Inference Engineer: You Already Know More Than You Think

If you know how to run applications on Kubernetes, you are already surprisingly close to understanding AI inference. Then someone says “tensor parallelism,” “KV cache,” or “prefill/decode disaggregation,” and suddenly it feels like you need another degree.

You don't.

This talk maps the world platform engineers already know to the inference stack they are being asked to operate. Scheduling becomes GPU scheduling. Resource requests become accelerator topology. Application replicas become hundreds of gigabytes of model weights. Autoscaling still exists, except cold starts can take minutes and state suddenly matters.

We will build an inference service from familiar Kubernetes primitives, then introduce the concepts that are actually new: models, tokens, accelerators, inference engines, batching, parallelism, routing, and the metrics that tell you whether any of it is working.

Come as a platform engineer. Leave speaking enough inference engineer to be dangerous.

Annie Talvasto

CNCF Ambassador & Sr. Manager at Upbound

New York City, New York, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top