Steve Green
Director, Slalom
Ann Arbor, Michigan, United States
Actions
Steve Green started coding back in 2003 and has never found a title that made him want to stop. He designed systems, founded and sold a consultancy, and has spent the years since leading engineering and delivery organizations.
He is a Director at Slalom, where he works with clients on cloud modernization and product engineering, usually on systems no one can afford to turn off. Before Slalom, he led delivery at Launch by NTT DATA, authored its enterprise AI governance operating model, and owned its production agentic platforms across multiple regulated environments. He holds an M.S. in Artificial Intelligence from Purdue, focused on management and policy.
Steve first spoke at KCDC in 2015 and has been part of the conference community ever since, presenting at CodeMash, IT/Dev Connections, CodeStock, DevNexus, and VS Live. He talks about engineering leadership, AI-assisted delivery, and the craft of building software together, which is to say he mostly talks about people, whatever the title on the program says. Between conferences, he mentors engineers and emerging managers, the same work in a smaller room. He lives outside Ann Arbor, Michigan.
Links
Area of Expertise
Topics
From Idea to Impact: A Practical Guide to Innovation
What if today’s "crazy" ideas are tomorrow’s game-changers? Inspired by Safi Bahcall's Loonshots, this session explores the power of disruptive innovation and how unconventional ideas can drive exponential growth. We’ll take a closer look at the hidden forces that influence innovation and show you how to nurture a culture where bold ideas are embraced, not ignored. You’ll leave with practical insights on creating an environment where breakthrough ideas flourish, helping your team build a sustainable pipeline of game-changing innovations.
Choas into Catalyst: Building Teams That Thrive on Change
Don’t just survive change—leverage it. Inspired by Nassim Nicholas Taleb’s concept of antifragility, this session shows you how to build teams that don’t just bounce back, but grow stronger in the face of disruption. Antifragile teams thrive on stress, uncertainty, and challenge—evolving and improving where others might break. We’ll dive straight into practical strategies for developing these high-performing teams, with a focus on leadership and team dynamics that drive real results. Learn how to foster adaptability, innovation, and a growth mindset—no matter what’s thrown your way. With actionable tools for building a culture of learning, rapid iteration, and psychological safety, you’ll leave ready to lead with distributed leadership and embrace diverse perspectives to create a team that thrives, adapts, and leads through uncertainty.
Ignite Your Team: Autonomy, Mastery, and Purpose
Why do some teams crush it while others fall short? This session explores the science of team performance by examining the essential human needs that fuel motivation and engagement. Drawing inspiration from Drive, Be Bad First, and An Elegant Puzzle, we will investigate the significant forces of autonomy, mastery, and purpose and how they contribute to a team’s success. Discover practical strategies for nurturing these elements within your team. Leave equipped with tools to evaluate your team’s current state, pinpoint growth opportunities, and implement changes that yield meaningful results.
Mastering Complexity
Software systems are getting more complex, but they don’t have to be a headache. This session gives you the tools to manage that complexity with smart software design. We’ll show you how to create loosely coupled systems that are easy to modify, extend, and evolve. You’ll learn how to manage change in decomposed systems and tackle incremental refactoring without the stress. Walk away knowing how to break down complex systems into manageable components that are more scalable, maintainable, and ready for whatever the future throws at them.
Writing Code for Humans
We have a lofty goal: programming style as documentation. Inspired by Steve McConnel’s “Code Complete”, Uncle Bob’s “Clean Code” and Andrew Hunt’s “The Pragmatic Programmer”, this session reviews best practices for writing code in a style that’s easy to create, maintain and understand. We’ll discuss concrete methods to get you there and give you a vocabulary for pragmatically evaluating code quality.
Various refactoring techniques, code smells, anti-patterns, and rules of thumb will be discussed, including fail fast, return early, separation of concerns, arrow code, magic numbers, the boy scout rule, being “stringly typed,” DRY, the step-down rule, table-driven methods, the importance of staying native, techniques for finding subtle redundancy, reinventing the square wheel, when to create a process, horizontal and vertical density, and simple design patterns. Within this session, we will refactor a confusing and ugly chunk of code into something beautiful, easy to read, and maintain. While examples are in C#, coders in any language should be able to follow along and apply the principles discussed. Though it seems like a lot, these topics will help developers rethink their approach to code.
Call the Shot: Plan-Do-Check-Act for AI Coding Agents
AI coding agents make generation faster while understanding, verification, and accountability remain human-speed. The mismatch appears when plans live in chat fragments, tests are written after the answer, and reviewers must reconstruct the reasoning from the diff. A team can save minutes during generation and spend them again as senior-engineer archaeology.
This session shows how to keep that from happening with a practical Plan-Do-Check-Act loop for coding agents. Through a prepared walkthrough, we approve a reviewable task graph before code exists, require the agent to call its shot before each test, compare the finished work with the plan, and turn what went wrong into a working agreement for the next cycle. You will see exactly where a person approves, interrupts, or redirects the work, and why each gate exists. If you're a developer or tech lead using coding agents, or deciding whether to, you'll leave with a repeatable pattern, concrete stop conditions, and a better way to explain why faster generation still needs deliberate judgment.
AI Is an Amplifier: Seven Capabilities That Shape Delivery
The same coding assistant can shorten feedback on one team and lengthen the review queue on another. Small batches, fast tests, and clear user outcomes turn generated code into learning. Weak version control, stale context, and large pull requests create more work nobody fully understands. Both teams can still report a successful rollout.
DORA's research helps explain why: AI amplifies the delivery system it enters. This session turns seven researched capabilities into signals a team can recognize in daily work, from growing pull requests and slow rollbacks to shadow use and disagreement about whether a change helped a user. We use a lightweight team assessment to reveal where engineers, leaders, and platform teams see the system differently, then choose the first condition worth improving. If you lead a team, or you're the engineer everyone asks whether the AI investment is working, you'll leave able to run that conversation with evidence and choose a practical next step before anyone buys another tool.
Design the Loop: Where Human Judgment Belongs in AI Workflows
A status update sounds low risk until leaders use it to decide what needs help, what can wait, and whether delivery is on track. When an agent compresses ten pages into five bullets, it is also deciding which uncertainty disappears. Human judgment has moved, even if the workflow still labels the output a draft.
This session follows one weekly status update from source material to the message leaders act on, watching what survives each handoff and what quietly falls away. We ask what the system can verify, what still requires judgment, and what happens when the reviewer never responds. The answers show where a human gate protects a real decision, when an agent should expose evidence before offering a conclusion, and how today's shortcut becomes tomorrow's process debt. If you lead a team or add AI to everyday work, you'll leave with a workflow map you can reuse the same week to decide where human judgment belongs.
From Doing to Supervising: Engineering Skill in the Age of Agents
Every engineer using a coding agent is becoming a supervisor of generated work, even if no one reports to them. When agents absorb routine fixes, developers get fewer repetitions that build debugging instincts, system knowledge, and the ability to recognize plausible wrong answers. They remain responsible for difficult exceptions even as the practice that prepared them grows quieter.
This is an old automation problem in a new setting: people inherit the difficult exceptions while getting less practice with the routine work that once prepared them. This session brings that research into everyday software development and makes the practice visible. We predict failures before generation, debug from symptoms with the agent off, and compare judgments with a team so system design, specification, and error recognition stay active. Whether you're early-career or newly responsible for agent output, you'll leave with a four-week practice plan that lets you use the tools without letting your judgment go quiet.
Govern the Capability, Not the Tool: Leading AI Adoption
Tool bans can reduce visibility faster than they reduce AI use. When the approved path cannot meet a real need and exceptions take weeks, engineers turn to personal accounts, copied prompts, and unofficial plugins to keep work moving. The demand remains, but the organization loses sight of the data, the capability, and the risk.
Governance works only when the safe path is fast enough for people to choose it. AI products change faster than most approval cycles, and the same capability soon appears under another name. Drawing on anonymized patterns from a large enterprise program, this session follows an AI use from local experiment to shared dependency and shows where tool-based approval loses the thread. We replace that thread with three durable questions about the capability, its data, and the consequence of failure, then connect the answers to proportionate review and visible ownership. If you lead engineers who are already using AI, you'll leave with a 30-day plan for making safe work easier to disclose, approve, and support.
Route, Evaluate, Repeat: Cost-Aware LLM Architecture in Production
A production LLM system can return correct answers while wasting money, repeating failed actions, and hiding the change from its dashboard. One real agent spent six hours retrying the same work through an expensive model while every visible health indicator stayed green. Quality, cost, and behavior need to be observable together because any one of them can conceal trouble in the other two.
This session reconstructs that failure and builds the architecture the dashboard was missing. We expose why a model was chosen, when a retry stopped being reasonable, and what evidence justified an escalation. From there, we add routing and evaluation in layers, starting with cheap programmatic checks and bringing in expert judgment where it can change the decision. If you build LLM features that must survive production traffic and financial scrutiny, you'll leave knowing when routing earns its complexity and the minimum event trail every agent should produce before something goes wrong.
Six Questions Before You Ship an Agent
Every agent has an authority boundary, whether the team designs it deliberately or discovers it during an incident. A support agent that can reset a password becomes a different risk when it can also change the account email. Useful permissions accumulate one at a time, and recovery can disappear before anyone asks how much autonomy the system should have.
This session turns the Cloud Security Alliance's autonomy levels into six questions a team can ask before an agent ships. We work those questions against the real account-takeover case, deciding which actions need approval, which boundaries belong in code, what must be reversible, and how the system should fail when no reviewer responds. Then we apply the same reasoning to prompt injection and untrusted content, where a carefully limited authority boundary can contain the damage even when a bad instruction gets through. If you build or approve agentic systems, you'll leave with one assessment you can take into a design review and clearer language for saying how far an agent may act on its own.
The Reversible Default: Architecture Decisions for Evolving Systems
Architecture decisions deserve different amounts of attention because the consequences of getting them wrong are not equal. A naming convention can absorb weeks of debate while a database choice slips through under calendar pressure. The useful questions are simple: what happens if we are wrong, and what will it cost to reverse the decision?
This session puts familiar architecture decisions on the table and asks how much ceremony each one has earned. We look at one-way and two-way doors, then use seams, contract boundaries, staged rollout, and sacrificial components to make expensive choices easier to undo. Lightweight architecture decision records preserve why a choice made sense, while fitness functions warn us when a temporary decision is quietly becoming permanent. Whether you're an architect, lead, or the developer whose pull-request comment just became a design meeting, you'll leave with a two-question risk test, an ADR template with explicit revisit triggers, and a faster route through decisions your team can safely change later.
Who's Accountable When the Agent Is Wrong?
Accountability remains organizational even when decision-making is automated. When ownership is vague, support points to engineering, engineering points to the vendor, and the vendor points to the model. Air Canada tried a version of that argument in 2024, but a British Columbia tribunal still held the company responsible for information delivered through its website.
This session uses the Air Canada ruling to pressure-test accountability before an incident assigns it by proximity. We ask who can explain an automated decision, who can stop the system, what evidence survives, and how a customer gets a correction. The exercise separates executive accountability, operational ownership, technical responsibility, and vendor obligations without allowing the answer to dissolve across them. If you deploy AI where customers or regulators can feel the result, you'll leave with a one-page ownership record and concrete answers to the three questions every incident asks: who approved this, who owns it now, and how would we have known?
Code Review Was Never Good at Finding Bugs
Code review is carrying more responsibility than it was designed to hold, and AI-generated volume is making the mismatch impossible to ignore. Teams ask senior engineers to inspect every generated line because experience feels like the safest control. Those same engineers also own the hardest design decisions, production problems, and mentoring work. When review becomes the answer to every defect, the queue grows while human attention gets thinner.
Code review still matters because it is where teams explain choices, share context, and decide whether a change belongs in the system. This session shows how to protect that human work by moving routine evidence earlier through test-first generation, static analysis, agent self-audit against an approved plan, and repository-level signals that tell us where to look. If you write, review, or merge code, you'll leave with a practical quality stack that puts routine checks in tools and reserves human attention for design, context, and ownership. You will also have a better answer than “review everything” when your team asks how to trust AI-assisted code.
Steve Green
Director, Slalom
Ann Arbor, Michigan, United States
Links
Actions
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top