Cuong Nguyen
Viettel, Cloud Solution Engineer
Hanoi, Vietnam
Actions
Cloud and Infrastructure Engineer with 6 years of experience in DevOps and distributed storage systems for Telco Cloud environments — from architecture design to day-2 operations and troubleshooting. Recently focused on integrating AI agents into Kubernetes observability to reduce manual toil in incident response, with a passion for sharing production-tested, reusable patterns with the open infrastructure community.
Area of Expertise
Topics
From 20 to 5 Person-Days: Standardizing OpenStack + Ceph Deployment
Standardizing OpenStack + Ceph deployment took us from 20 person-days per cluster with two engineers to 5 with one. Deployment time was the smaller half of the payoff.
Environment: Viettel private cloud — 10 OpenStack clusters with Ceph, bare metal, each deployed for a separate tenant organization across different data centers. Standardization was non-optional: every difference tolerated at deployment became one we had to remember at incident time, for a cluster we might not touch again for months.
Where the 20 person-days went — five recurring sources:
— Non-standardized tooling
— Manual hardening: many-item checklist applied by hand
— Manual integration: OpenStack,Ceph, monitoring, logging
— Piecemeal validation: no end-to-end gate, defects surfaced at handover
— Environment drift: failures no runbook anticipated
Drift produces defects, piecemeal validation finds them too late, and hand-built clusters keep charging in incidents and support long after handover. The highest-leverage fix: gate the environment before deploying, not the deployment after.
Attendees leave with a five-category diagnostic framework and a pre-deployment validation checklist for bare-metal OpenStack + Ceph.
Investigate First, Decide Second: The Missing Step in Kubernetes Alert Response
"We automated the investigation, not the fix."
Every DevOps engineer knows the drill. Alert fires. Open AlertManager, switch to Kibana, check Grafana. Correlate manually across three tools. At 2AM, this takes a hour before you know enough to act.
On Kubernetes 1.33, we integrated an AI agent that reads alerts from AlertManager, queries logs from Elasticsearch, pulls metrics from Prometheus assembling a structured investigation before anyone is paged. The engineer receives a brief, not a fire alarm.
Three scenarios, same stack, different outcomes:
- CrashLoopBackOff: agent identifies OOMKill pattern across 3 restart cycles — engineer approves fix in 5 minutes, not 1 hour
- ImagePullBackOff: Kibana is empty because container never started
- OOMKilled: Prometheus memory trend reveals misconfigured resource limit — fix proposed before engineer opens a single tab
Attendees leave with a reusable investigation pipeline and approval gate design for AI-assisted alert response on Kubernetes.
Changing the Engine Mid-Flight: Zero-Downtime Ceph Upgrades
"Upgrade Ceph in our telco cloud with zero downtime." That is a mandatory requirement every 1–2 years to ensure security patches, bug fixes, and continued support from vendors and the community. We have a Ceph Cluster version with 10–50 nodes deployed on bare-metal running 5G Core workloads (AMF, UPF,...).
This talk covers a first-hand upgrade with three real failure scenarios:
- Monitor quorum loss: root cause analysis, recovery sequence, and how to prevent quorum degradation during daemon rolling restarts
- OSD storms triggered by rebalancing that threatened cluster stability
- Incompatible client versions silently blocking the upgrade path
Beyond failure recovery, we'll share the upgrade sequencing strategy we developed — covering pre-flight checks, daemon upgrade ordering (MGR → MON → OSD → MDS/RGW).
Attendees leave with a reusable pre-upgrade checklist and sequencing framework for bare-metal Ceph in high-stakes environments.
KubeCon + CloudNativeCon Japan 2026 Sessionize Event
KCD & OpenInfra Days Vietnam 2026 Sessionize Event
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top