Session

The Agent That Takes the First Ten Minutes of an Incident

Every incident starts the same way: someone gets paged, opens six tabs, and spends ten minutes assembling context before anyone can make a decision. That assembly work is well-shaped for an agent: bounded, repetitive, read-mostly - and it is where autonomous operations should start, rather than with the fantasy of an agent that fixes production by itself.
We build and run a triage agent on Azure that wakes on an alert, pulls the signal it needs from Azure Monitor and Log Analytics, correlates it against recent deployments and PRs, and produces a written incident brief with a ranked hypothesis and a citation for every claim it makes. Then it opens a draft PR for the one remediation we are willing to let it propose — and stops there, on purpose.
The second half is the part that makes this operable: what the agent is allowed to touch and how that boundary is enforced with a workload identity rather than a prompt instruction, what the blast radius looks like when the agent is wrong, how to evaluate a triage agent when every incident is a sample size of one, and the honest failure mode — an agent that writes a confident brief about the wrong service and sends three engineers in the wrong direction.
Attendees leave with the read-only-first pattern, the identity and approval boundaries worth setting before the first deploy, and a clear view of which parts of on-call should stay human.

Sarang Brahme

Tech Leader, Code and Curiosity

Vancouver, Canada

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top