Session
SWE-in-a-team: a coding benchmark for software factories
Coding benchmarks usually grade the destination: a function, a patch, a passing environment. Engineering teams do not work that way. We take a ticket through planning, CI, review, deployment and testing, then return to the code when one of those gates objects.
I built SWE-in-a-team to measure that unit of work. Thirteen coding agent and model configurations each worked through the same 20 tickets in a private, real-world-like full-stack SaaS app. That produced 260 missions across bugs, security fixes, migrations, integrations, performance work and features. Every builder worked with the same planner, reviewer and tester. A ticket counted as resolved only when the delivery loop accepted it and hidden tests agreed.
The results challenged my own reflex to reach for the strongest model. A budget open-weight builder was sent back 13 times, corrected course and resolved all 20 tickets for $3.34 each. A frontier builder sailed through nearly every internal gate, cost almost twice as much and still failed a hidden test. The harness alone changed cost by 22% to 41% for the same models.
In this talk, I will show where the money and time went, what CI, review and testing caught, and why convergence under correction can matter more than a perfect first draft. You will leave with a practical way to evaluate coding agents inside the delivery system they must actually survive.
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top