Session
Unlocking Liquid Compute: Scaling AI & LLM Inference via Kubernetes Dynamic Resource Allocation
As generative AI and Large Language Models (LLMs) transition from experimental batch jobs to business-critical production infrastructure, orchestrating expensive hardware like GPUs/TPUs has become an operational hurdle. The traditional Device Plugin framework often falls short, forcing platform engineers into manual node pinning, restrictive "all-or-nothing" allocations, and static node selectors.
Enter Dynamic Resource Allocation (DRA)—the game-changing evolution in Kubernetes device management.
In this technical session, we will break down how DRA shifts the paradigm from rigid hardware assignments to a highly fluid, "liquid" resource pool. We will explore how DRA decouples hardware inventory from workload requirements using ResourceSlices and ResourceClaims, allowing the Kube-scheduler to make topology-aware, capability-based scheduling decisions.
Attendees will learn production-ready best practices for running high-volume, low-latency LLM inference workloads. We will cover:
- How to define custom DeviceClasses as abstract blueprints for developers.
- Strategies for fine-grained resource sharing (like VRAM requirements and interconnect topology).
- Essential cluster administration safeguards, including DRA driver deployment, liveness probing, and graceful node draining to minimize scheduling delays.
Whether you are scaling open-source foundational models or building internal developer platforms for data science teams, this talk will give you the architectural blueprint to optimize accelerator utilization, eliminate resource drift, and slash your AI infrastructure TCO.
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top