Session

Engineering Moneyball: Proving ROI on Tooling Spend

A customer told me last month he feels like Billy Beane. He can see his team's good. He can't get a number to say it.

This session is the model behind that feeling. Every merged commit is scored deterministically into one throughput unit, with token and tooling spend divided by it; each commit is tied to the roadmap. Then where the model breaks, which gets the most stage time: a customer team running 45% of its work with no roadmap link, and orgs burning tokens with roadmap delivery flat.

Oakland A's, 2002. About $40M of payroll against the Yankees' $126M. Beane couldn't outspend them, so he stopped measuring what they measured. Every club had rooms full of data. Beane found the number that predicted winning and bet on it while the room laughed. On-base percentage. The A's finished 103-59, one more loss than the Yankees, on a third of the money.

Ask an engineering org how it measures delivery, and you'll hear velocity, story points, commit counts, and a leader who used to ship and now trusts his gut. That's the eye test. As a former CTO, I was that leader, and the numbers I reported upward were the numbers I'd been handed.

Velocity tells you a team's busy. It runs on self-estimates, so the team grades its own homework and reports the grade as a metric, and it doesn't compare across teams or, often, across sprints inside one team. AI coding tools finished off the estimate. When an engineer ships in an afternoon what used to take three days, the number was fiction before the sprint started. The unit changed,d and nobody recalibrated it.

The stat this market underprices is roadmap delivery. Did the roadmap move, how fast, and is that speed trending? It's harder to measure than counting commits, which is why people count commits.

Here's the method. Each merged commit is scored on the depth of work it represents: how the change classifies, how much of the codebase it affects, where it sits in the architecture, and whether it corrects a defect someone else introduced. The sum is one unit, Engineering Throughput Value. We split the work into growth, maintenance, and fixes, then tie each commit to a Jira or Linear item, so "on the roadmap" is a link rather than an opinion. Then we divide tooling and token spend by throughput shipped. Beane's whole game was cost per win.

One customer team, real numbers: 80% faster than a year ago, $25k a month in tooling, about $200 per ETV shipped, roadmap delivery 36% faster, and 45% of the work carrying no roadmap link at all. Our own published healthy line is 75% linked, so that team's drifting. Why it's drifting is the part worth an hour.

Then the failure modes, in detail. Where file-level classification stands in for deeper structural analysis. What a rewritten repo history does to the score. What squash policy and monorepo layout do to the number without touching the work. Everything a commit can't see, starting with the non-coding overhead that fills most of a developer's week. And the case that never makes a marketing page: high token burn with roadmap delivery flat, which we see often.

Team-level throughput, never a ranking of individuals. Code stays on your own infrastructure. No demo.

We have a commercial interest in this model, so the scoring design, sample boundaries, and known limits are published at research.navigara.com, and the live index at 500.navigara.com runs the same model across VS Code, React, TypeScript, Next.js, and 62 other repositories. You can check my math while I'm still on stage.

Jirka Bachel

Co-Founder & CEO

San Francisco, California, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top