Session
The Future of AI Inference is Small, Distributed, and Budget-Conscious
By 2030, 90% of AI workloads will be inference—mostly smaller models at smaller organizations. Yet KubeCon talks focus on massive GPU clusters for frontier labs. We're building infrastructure for the wrong market.
This lightning talk argues for a shift: make AI inference accessible to budget-conscious organizations using CNCF projects we already have. I'll sketch a vision combining Cluster API (elastic infrastructure), HAMi (GPU sharing), and Kaito (optimized serving) to create cost-effective, distributed inference for the 90% of organizations who can't afford hyperscaler patterns.
Let's stop optimizing for OpenAI and Anthropic, and start optimizing for everyone else.
Mohamed Belgaied Hassine
Consulting Architect at Mirantis
Munich, Germany
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top