Session

The Future of AI Inference is Small, Distributed, and Budget-Conscious

By 2030, 90% of AI workloads will be inference—mostly smaller models at smaller organizations. Yet KubeCon talks focus on massive GPU clusters for frontier labs. We're building infrastructure for the wrong market.

This lightning talk argues for a shift: make AI inference accessible to budget-conscious organizations using CNCF projects we already have. I'll sketch a vision combining Cluster API (elastic infrastructure), HAMi (GPU sharing), and Kaito (optimized serving) to create cost-effective, distributed inference for the 90% of organizations who can't afford hyperscaler patterns.

Let's stop optimizing for OpenAI and Anthropic, and start optimizing for everyone else.

Mohamed Belgaied Hassine

Consulting Architect at Mirantis

Munich, Germany

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top