Session
LLM Inference on Edge Kubernetes: Patterns for Offline, Constrained, and Heterogeneous Devices
Edge LLM inference isn’t “cloud inference, but smaller.” Edge deployments must handle intermittent connectivity, tight memory/compute budgets, and heterogeneous accelerators (CPU-only, iGPU, small GPUs, NPUs). Teams that reuse cloud assumptions (always-on control plane, uniform hardware, centralized logging) hit failure modes that look like randomness: cold starts that take minutes, unpredictable latency, silent model corruption, and observability gaps when devices go offline. This talk presents Kubernetes-native patterns for shipping and operating LLM inference on edge clusters (single-node and small multi-node): model artifact packaging and rollouts (versioning, integrity checks, rollback safety), scheduling/isolation across heterogeneous nodes, offline-first operations (local queues + eventual sync), and observability that still works with delayed export (local buffering + minimal health signals).
Ref Repo: https://github.com/harshuljain13/llm-inference-at-scale
Harshul Jain
Audible, Senior Software Engineer
Newark, New Jersey, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top