Session

Azure Node Management: Surviving a PDB Deadlock During Routine Upgrades

It was supposed to be a standard Tuesday. We were rotating a TLS secret for an Ingress in our elastic namespace on AKS routine stuff. But within minutes, what should've been a quick config change turned into a complete deadlock.

Our Pod Disruption Budgets did exactly what they were designed to do: block evictions. The problem? They worked too well. They froze the upgrade path, trapped our pods, and prevented the Azure nodes from scaling down. The cluster couldn't heal itself, and we were stuck.

In this session, I'll walk through what went wrong and how we manually recovered the namespace without losing data. We'll go beyond basic kubectl commands and cover:

The Trap: How a mismatch between PDBs and deployment rolling updates creates a dependency loop that locks everything up.
The Fix: The actual cordon, drain, and uncordon steps we used on AKS nodes to force a reset. I'll show you what worked (and what didn't).
The Lesson: How to calculate PDB values that won't break your cluster during Azure's upgrade cycles so you don't end up troubleshooting at 3am like we did.

Naman Kaley

Docker Captain | Docker Certified Associate | Hands-On Transformative AI Leader | Architect of Generative AI & Neuroscience-Inspired Systems | Solutions Architect

Jaipur, India

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top