Session

Multi-Agent Systems in Production: Build It, Evaluate It, Gate It, Ship It

Everyone can build an agent. Shipping one you can trust is the hard part - agents never behave the same way twice, so "looks good on my machine" doesn't survive production.

This hands-on workshop is for engineers past the demo phase. We treat a multi-agent system like real software: something you evaluate, gate, and ship on a pipeline - not something you eyeball and hope.

We build a multi-agent system, then put it on rails:

- Evaluation for non-deterministic output: how to measure "good" when the answer changes every run, and build eval sets you actually trust.
- Gate-keeping: the checks an agent must pass before it's allowed to act or ship, and where human-in-the-loop belongs.
- Frameworks and tooling: what to actually reach for - orchestration, eval harnesses, and observability.
- CI/CD for agents: wiring evals and gates into a pipeline so regressions get caught automatically, not in prod.
- Red teaming as a gate: prompt-injection and goal-hijack checks that run inside the pipeline, one gate among many.

You leave with a working system that evaluates itself, blocks its own bad releases, and ships on a pipeline you understand.

Daniel Ostrovsky

AI Architect at Payoneer | Full Cycle Development Expert | Public Speaker | Open Source Contributor |

Tel Aviv, Israel

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top