Session
The Ephemeral Scale: Managing 10k K8s Clusters a Day
Kubernetes is notoriously difficult to learn because "Day 2" operations require breaking things - something you can’t easily do in a production environment. To solve this, we built a platform at KodeKloud that spins up thousands of fully isolated, short-lived Kubernetes clusters every hour for students to experiment with.
In this technical deep dive, I will share how we leverage the CNCF ecosystem to build an automated learning engine:
1. The Orchestrator: How we use K3s and custom controllers to manage the lifecycle of hyper-ephemeral clusters with sub-30-second cold starts.
2. Networking and Isolation: Using Cilium or Calico to ensure that thousands of students can run "root" workloads in isolated sandboxes without cross-contamination.
3. How we monitor the health of thousands of clusters that only exist for 60 minutes using Prometheus and Grafana.
4. How we implemented automated reconciliation to ensure a student can "reset" their environment to a known good state instantly.
This session is for anyone interested in Platform Engineering, Education Technology, or those looking for advanced patterns in ephemeral infrastructure and resource multi-tenancy.
Abhinav Sharma
Site Reliability Engineer at KodeKloud | Microsoft MVP | GSOC @OpenSUSE | GitHub Campus Expert
Jaipur, India
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top