Session

Measuring Developer Productivity After AI: Year Two

Two years ago we mandated GitHub Copilot and Claude Code for 600 engineers and asked the obvious question: how do we measure whether this is working?

The early answer was that it was working very well. Deploy frequency rose 30% and developer satisfaction rose 14%, and for a while those were the numbers we reported. What we did not see for months was everything moving the other way underneath them. Code quality metrics, such as cyclomatic complexity, maintainability index, test coverage, and duplication, deteriorated by 15%, and bugs and performance issues followed. Platform costs climbed 11%, driven by infrastructure components the agents introduced and by inefficient use of what we already had, both of which a timely challenge to the agents would have caught.

The uncomfortable part is that our DevOps Research and Assessment (DORA) metrics were not lying to us. Change failure rate and lead time registered the damage. We were reading them monthly while agents were shipping daily.

So we changed the workflow rather than the dashboard. We stopped letting agents run unchecked and required a written plan, verified by a human, before any change was made. The gate itself is the obvious part; where it sits and what it costs are not. Code quality recovered to where it was before we adopted AI at all, and the gains held. Velocity and satisfaction stayed up. Platform costs stabilized. Developer tooling costs did not, and are now 20% higher than when we started, and still climbing. That is the trade-off we made, and this talk is about whether it was worth it.

We'll cover: (1) which signals moved first, which lagged by months, and why velocity is the last one to warn you; (2) the plan and verify checkpoint in detail, including where it sits, who owns it, and what it costs in developer time; (3) what the recovery actually looked like, and how long each signal took to come back.

You'll leave able to instrument your own AI rollout so quality regressions surface in weeks rather than quarters, with a checkpoint design you can put in front of your teams next sprint.


OUTLINE (45 minutes including Q&A)

1. The numbers we reported (4 min). Deploy frequency +30%, satisfaction +14%, two years, 600 engineers, Copilot and Claude Code. Presented as they were reported internally at the time. Let the audience agree the rollout was a success.

2. The numbers we weren't looking at (7 min). Quality down 15%. Platform costs up 11%, split between infrastructure the agents introduced and inefficient use of existing infrastructure. Bugs and performance issues. Both sets of curves on one timeline.

3. Which signal moved when, and what else could explain it (6 min). The lag order, and the case that velocity is a trailing indicator of nothing useful. Then the confounders, stated before the audience raises them: the tools improved over the window, teams changed, the codebase aged. What the data can and cannot prove.

4. Why DORA wasn't the problem (5 min). Change failure rate and lead time registered the damage. The failure was cadence: monthly reads against daily shipping. The argument against the popular "AI broke our metrics" position.

5. Two changes, not one (12 min). First, the measurement change: what moved to a sprint cadence and which signals were worth watching that often. Second, the plan and verify checkpoint: what triggers it, what a plan must contain, who verifies, what happens on rejection, what it costs in developer time, and what still runs unchecked. Concrete enough to copy.

6. What recovery looked like (5 min). Quality back to pre-AI levels within 2 to 3 sprints. Velocity and satisfaction held. Platform costs stabilized. Tooling costs still climbing at +20%. Close on the trade-off honestly.

7. Q&A (6 min).

30-minute version: merge 3 into 2, cut 4 to 3 min, 5 to 9 min, Q&A to 5.
60-minute version: +5 to section 5 for a worked example of one plan and its verification, +4 to section 6 for the cost model, +2 for what we would do differently today, Q&A to 10.

TAKEAWAYS (3)

1. Match your metric review cadence to your agents' shipping cadence, and diagnose whether a monthly or quarterly read is hiding a quality decline you are already paying for.

2. Design a plan and verify checkpoint for agent-driven changes, including where it sits in the workflow, who owns verification, and what it costs in developer time.

3. Estimate the true cost of agentic development for your organization, separating platform spend, which stabilizes, from tooling spend, which does not.

Uroš Miletić

IPS, Chief Technology Officer

Prague, Czechia

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top