Session
Azure Node Management: Surviving a PDB Deadlock During Routine Upgrades
It was supposed to be a standard Tuesday. We were rotating a TLS secret for an Ingress in our elastic namespace on AKS routine stuff. But within minutes, what should've been a quick config change turned into a complete deadlock.
Our Pod Disruption Budgets did exactly what they were designed to do: block evictions. The problem? They worked too well. They froze the upgrade path, trapped our pods, and prevented the Azure nodes from scaling down. The cluster couldn't heal itself, and we were stuck.
In this session, I'll walk through what went wrong and how we manually recovered the namespace without losing data. We'll go beyond basic kubectl commands and cover:
The Trap: How a mismatch between PDBs and deployment rolling updates creates a dependency loop that locks everything up.
The Fix: The actual cordon, drain, and uncordon steps we used on AKS nodes to force a reset. I'll show you what worked (and what didn't).
The Lesson: How to calculate PDB values that won't break your cluster during Azure's upgrade cycles so you don't end up troubleshooting at 3am like we did.
Naman Kaley
Docker Captain | Docker Certified Associate | Hands-On Transformative AI Leader | Architect of Generative AI & Neuroscience-Inspired Systems | Solutions Architect
Jaipur, India
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top