Uroš Miletić
IPS, Chief Technology Officer
Prague, Czechia
Actions
Uroš Miletić is Chief Technology Officer at IPS, responsible for architecture, software delivery, and security across an engineering organization of several hundred developers. He has spent his career building platforms on Microsoft Azure and now works mostly on AI adoption, and on the problems that come with it at scale. He is interested in what only becomes visible after the pilots are over.
Area of Expertise
Topics
AI Security for a safer tomorrow
As AI becomes increasingly integrated into our daily lives, ensuring its security is paramount. This talk delves into the critical importance of safeguarding AI systems against emerging threats, such as prompt injection, sensitive information exfiltration, data poisoning, etc. Join to discover how cutting-edge research and cloud-based technologies are shaping a safer tomorrow for AI technologies.
This is a level 200 talk for AI enthusiasts and security specialist. Although some knowledge of GenAI and cybersecurity concepts is welcome, it is not required to enjoy the session.
ChatGPT under the hood
Ever wanted to peek behind the curtain of ChatGPT? It might have been a single LLM answering your prompts in its inception, but today it is a rather complex system using state-of-the art GenAI technologies.
In this session, we'll take you on a journey through the overall architecture of agentic systems and how they power applications like ChatGPT. We'll uncover the key components and design principles that make these systems tick. From orchestrating natural language processing capabilities to integrating various subsystems, this talk will provide a comprehensive overview of the techniques and best practices for building advanced conversational agents.
This talk is rated at level 300. It is perfect for AI practitioners and developers with a solid grasp of software architecture concepts and LLMs, who are eager to deepen their understanding of agents and agentic systems. The concepts presented are abstract and not dependent on any specific technology, making this presentation interesting for non-technical professionals as well.
Microsoft Copilot under the hood
Ever wanted to peek behind the curtain of Microsoft Copilot? It might have been a single LLM answering your prompts in its inception, but today it is a rather complex system using state-of-the art GenAI technologies.
In this session, we'll take you on a journey through the overall architecture of agentic systems and how they power applications like Microsoft Copilot. We'll uncover the key components and design principles that make these systems tick. From orchestrating natural language processing capabilities to integrating various subsystems, this talk will provide a comprehensive overview of the techniques and best practices for building advanced conversational agents.
This talk is rated at level 300. It is perfect for AI practitioners and developers with a solid grasp of software architecture concepts and LLMs, who are eager to deepen their understanding of agents and agentic systems. The concepts presented are abstract and not dependent on any specific technology, making this presentation interesting for non-technical professionals as well.
Common pitfalls in building generative AI platforms
Curious about the challenges that can arise when creating generative AI solutions? While such technologies offer incredible potential, they also come with their own unique set of hurdles.
In this session, we'll explore the common pitfalls encountered during the development of generative AI platforms. We will discuss the appropriate contexts for using generative AI solutions, how to avoid jumping into complex solutions prematurely, and why keeping human in the loop is essential for success. By examining real-world examples, we'll highlight the mistakes to avoid and the best practices to follow to ensure your generative AI projects have successful outcome.
This talk is rated at level 200. It is ideal for AI practitioners and developers with a solid understanding of AI concepts. The insights shared will be valuable for both technical and non-technical professionals interested in the practical aspects of generative AI development.
Common pitfalls in AI adoption
Curious about the challenges of integrating AI into your organization? While AI technologies offer immense potential, they also present unique hurdles.
In this session, we will explore the common pitfalls businesses face when adopting AI. We will discuss the right contexts for implementing AI solutions, how to avoid diving into complex projects too soon, and the importance of maintaining human oversight. Through real-world examples, we will highlight mistakes to avoid and best practices to ensure your AI initiatives succeed.
This talk is ideal for business leaders, decision-makers, and professionals interested in the practical aspects of AI adoption. The insights shared will benefit both technical and non-technical audiences navigating the complexities of AI integration.
How do CTOs use Generative AI?
By now, it’s clear that Generative AI is reshaping every stage of the software development lifecycle. But what does that mean for CTOs—the people responsible for driving technical excellence and fostering innovation?
My time as a CTO happened to align with the rise of Generative AI. I embraced these unexplored tools, and in turn, they helped me automate repetitive tasks and boosted the speed and quality of my work. In this session, I’ll share how Generative AI is changing the way tech leaders operate, from decision-making, effort estimation, to rapid prototyping. We’ll dive deeper on concepts like AI agents, vibe coding, and multimodal models.
This talk is rated at level 200. While we’ll get into a lot technological concepts, the talk is designed to be accessible to a broader audience.
Build Agentic Group Chat using Semantic Kernel and SignalR
Ever felt that traditional AI chats were a one-way street? Modern AI platforms are evolving from single-turn, one-sided interactions into collaborative, multi-agent dialogues where users are active participants. These systems don’t just respond; they reason, coordinate, and engage, creating dynamic conversations that feel more like teamwork than Q&A.
In this session, we’ll explore how to design and implement an Agentic Group Chat system that includes the user as an active participant in the conversation. You’ll learn how Semantic Kernel enables orchestration of multiple AI agents and how SignalR powers real-time user interactions. We’ll demonstrate these concepts using an open-source project focused on development effort estimation, showcasing how agentic systems can collaborate with humans to solve complex problems.
This talk is rated Level 300 and is ideal for developers, AI practitioners, and architects who have a solid understanding of .NET AI stack.
Confidently incorrect
By now, it’s clear that LLMs are transforming how we interact with information, automate workflows, and build intelligent applications. But what happens when these powerful systems “hallucinate”, i.e., confidently generate answers that sound right but are actually wrong?
In this session, we’ll demystify the phenomenon of AI hallucinations: why they happen, how often they occur in today’s top models, and why they matter for everyone from developers to business leaders. Drawing on the latest research and real-world examples, we’ll explore the technical roots of hallucinations, their impact on trust, safety, and operational efficiency. We will also take a look at evolving toolkit for reducing hallucinations—from retrieval-augmented generation and prompt engineering to automated fact-checking and human feedback.
This talk is rated at level 300. While we’ll dive deep into the technical causes and mitigation strategies for hallucinations, the session is designed to be accessible to a broad audience, including technologists, leaders, and anyone curious about the future of trustworthy AI.
Bursting the AI bubble
AI has become the centerpiece of innovation and investment, promising transformative change across all industries. But behind the hype lie critical challenges, both technological and economic, that threaten its long-term viability. What happens when scaling laws collide with physical limits, infrastructure costs outpace ROI, and organizations discover that AI’s promised gains aren’t materializing?
In this session, we’ll examine the cracks in today’s AI landscape, from the unit‑economics that make widespread adoption challenging and the unrealistic expectations shaping investment decisions. Drawing on recent research and long-term key metrics from large scale AI deployments, we’ll explore why the current trajectory is unsustainable and outline what must structurally change for AI to deliver long-lasting impact.
This talk is rated at level 200 and is designed for both business leaders and technologists, and anyone curious about the future of the AI industry.
Bringing Biztalk Maps to Azure
Maps often hold the most business know-how and represent the heart of any integration effort. They are, however, notoriously hard to modernize due to tight coupling with the integration platform. This is exactly the challenge our team faced during the migration of an enterprise BizTalk platform to Azure Integration Services. Despite vendor assurances that our existing toolkit could be fully reused, reality told a different story. That forced us to reinvent how we create and maintain maps, using a modern approach designed to evolve as surrounding technologies change.
In this session, we will explore a complex migration story where we moved critical transformations from a BizTalk environment and made them usable in Azure Logic Apps. We will dive into an alternative framework we developed to build, refine, and validate mapping logic, built entirely on Azure Integration Services and enabled by the power of generative AI.
This talk is rated at level 300. It is ideal for integration practitioners and stakeholders looking for a modern twist on decades-old mapping concepts. The insights shared will be valuable for both technical and non-technical professionals interested in the practical realities of BizTalk Mapper and Azure Integration Services.
We cut token usage by 70%, but quality went up
Over the last two years, AI coding assistants stopped being a nice-to-have and became the way our teams work. Then the vendors moved to usage-based billing, and a previously predictable cost started behaving like a rollercoaster. Projected across several hundred developers, we were looking at a 4-5x increase in our AI tooling bill.
Both obvious alternatives were bad. Absorb the cost, and finance would own the engineering AI roadmap. Restrict usage, and we would throw away the momentum our teams so meticulously built. So we looked further for what was actually driving the token usage, expecting a billing problem.
What we found was a working-habits issue. On an identical feature, an undisciplined session burns most of its budget on reasoning tokens the developer never sees. The most expensive sessions were the ones using the largest model on the smallest questions, with contexts growing without restrictions, and developers prompting their way toward a solution instead of driving one. Fixing these habits cut measured token usage by roughly 70% on the same work. Teams doing it reported better output and better work satisfaction, with velocity unchanged.
We'll cover: (1) the three levers that actually drive spend: model choice, conversation size, and interaction count; (2) two changes you can make the same afternoon: matching model capability to task and rewriting your agent instructions; (3) spec-driven development, the workflow change behind most of the savings.
Our developers resisted at first, and they were right to: this asks them to carefully choose the model, keep context tight, and verify output at every step. Nobody got slower, because a disciplined session needs less back and forth than an undisciplined one. The friction was the point.
You'll leave able to estimate your own token spend, identify which habits burn the most for the least return, and run a spec-driven session the next morning.
OUTLINE (45 minutes, including Q&A)
0-3 — The rug pull. Usage-based billing arrives with no meaningful notice. Projected 4-5x across several hundred developers. Two bad options.
3-6 — Shared vocabulary and the three cost dimensions. Context, token, agent, reasoning, output. Establishing that reasoning tokens are billable and invisible.
6-13 — Lever 1: model choice. The 0.33x to 15x spread. What large models are genuinely needed for versus what small models handle fine.
13-23 — Lever 2: conversation size. Input and output both bill. Instruction files as prevention; output compression; the instruction rewrite from 2,500 tokens to 300.
23-33 — Lever 3: interaction count. Vibe coding versus spec-driven on an identical feature, with token comparison and live demo (5-6 min, recorded fallback).
33-37 — What it cost us and what we got. Developer resistance, the 70% result, honest scope of what was and wasn't measured. Frameworks to start with. Close on the friction being the point.
37-45 — Q&A.
TAKEAWAYS (3)
Estimate what your organization's AI assistant usage costs per developer per month, and identify which of the three levers is driving the largest share.
Choose between large and small models per task type, using capability boundaries that hold in practice rather than defaulting to the most capable option.
Run a spec-driven session on a real feature and compare its token consumption against the equivalent unplanned session.
Seamless AI Software Development Life Cycle
What if your software delivery process could move smoothly from idea to production with AI assisting at every step? Not just code completion, but a connected, intelligent flow across the entire development life cycle.
In this session, we will explore how modern AI capabilities can transform the way teams gather requirements, shape solution designs, generate and review code, and operate software in production. Instead of focusing on isolated prompts or code completion, we will look at how agentic workflows can be connected into a more seamless engineering process. The session will include practical showcases of building and running such frameworks in our organization. We will also discuss what works well, what does not, and where human judgment is still absolutely necessary. Topics such as quality gates, traceability, review workflows, and keeping developers in control will be covered from a practical engineering perspective.
This talk is rated at level 300. It is designed for software engineers, architects, and technology leaders who want to see how AI can be applied beyond coding productivity gains. Participants will leave with a clearer view on how to build agentic SDLC workflows that are ready for real-world software delivery.
Sovereign AI by Design
Ever tried building cool new AI feature, only to discover that the required data cannot leave a regulated boundary? Many architectures assume that data can be sent to an AI model with minimal friction. In real organizations, things are rarely that simple.
In this session, we’ll look at practical ways to design AI agents when data sovereignty is a core architectural constraint. We’ll cover patterns for handling confidential data, setting retrieval boundaries, and enforcing compliance, grounded in real challenges we face daily. What happens when a manufacturing client wants to use generative AI to improve operational decision-making but cannot expose sensitive factory floor data? Or when a financial institution wants to automate customer support without disclosing confidential client information to AI models? How do you deliver AI value while keeping critical organizational data protected?
This talk is rated at level 300. It is intended for architects and technical leaders who are building AI platforms in enterprise or regulated environments. Attendees should have a basic understanding of LLM-based applications, retrieval-augmented generation, and modern software architecture, but the session will remain practical and technology-neutral enough to be valuable for product and business stakeholders as well.
Measuring Developer Productivity After AI: Year Two
Two years ago we mandated GitHub Copilot and Claude Code for 600 engineers and asked the obvious question: how do we measure whether this is working?
The early answer was that it was working very well. Deploy frequency rose 30% and developer satisfaction rose 14%, and for a while those were the numbers we reported. What we did not see for months was everything moving the other way underneath them. Code quality metrics, such as cyclomatic complexity, maintainability index, test coverage, and duplication, deteriorated by 15%, and bugs and performance issues followed. Platform costs climbed 11%, driven by infrastructure components the agents introduced and by inefficient use of what we already had, both of which a timely challenge to the agents would have caught.
The uncomfortable part is that our DevOps Research and Assessment (DORA) metrics were not lying to us. Change failure rate and lead time registered the damage. We were reading them monthly while agents were shipping daily.
So we changed the workflow rather than the dashboard. We stopped letting agents run unchecked and required a written plan, verified by a human, before any change was made. The gate itself is the obvious part; where it sits and what it costs are not. Code quality recovered to where it was before we adopted AI at all, and the gains held. Velocity and satisfaction stayed up. Platform costs stabilized. Developer tooling costs did not, and are now 20% higher than when we started, and still climbing. That is the trade-off we made, and this talk is about whether it was worth it.
We'll cover: (1) which signals moved first, which lagged by months, and why velocity is the last one to warn you; (2) the plan and verify checkpoint in detail, including where it sits, who owns it, and what it costs in developer time; (3) what the recovery actually looked like, and how long each signal took to come back.
You'll leave able to instrument your own AI rollout so quality regressions surface in weeks rather than quarters, with a checkpoint design you can put in front of your teams next sprint.
OUTLINE (45 minutes including Q&A)
1. The numbers we reported (4 min). Deploy frequency +30%, satisfaction +14%, two years, 600 engineers, Copilot and Claude Code. Presented as they were reported internally at the time. Let the audience agree the rollout was a success.
2. The numbers we weren't looking at (7 min). Quality down 15%. Platform costs up 11%, split between infrastructure the agents introduced and inefficient use of existing infrastructure. Bugs and performance issues. Both sets of curves on one timeline.
3. Which signal moved when, and what else could explain it (6 min). The lag order, and the case that velocity is a trailing indicator of nothing useful. Then the confounders, stated before the audience raises them: the tools improved over the window, teams changed, the codebase aged. What the data can and cannot prove.
4. Why DORA wasn't the problem (5 min). Change failure rate and lead time registered the damage. The failure was cadence: monthly reads against daily shipping. The argument against the popular "AI broke our metrics" position.
5. Two changes, not one (12 min). First, the measurement change: what moved to a sprint cadence and which signals were worth watching that often. Second, the plan and verify checkpoint: what triggers it, what a plan must contain, who verifies, what happens on rejection, what it costs in developer time, and what still runs unchecked. Concrete enough to copy.
6. What recovery looked like (5 min). Quality back to pre-AI levels within 2 to 3 sprints. Velocity and satisfaction held. Platform costs stabilized. Tooling costs still climbing at +20%. Close on the trade-off honestly.
7. Q&A (6 min).
30-minute version: merge 3 into 2, cut 4 to 3 min, 5 to 9 min, Q&A to 5.
60-minute version: +5 to section 5 for a worked example of one plan and its verification, +4 to section 6 for the cost model, +2 for what we would do differently today, Q&A to 10.
TAKEAWAYS (3)
1. Match your metric review cadence to your agents' shipping cadence, and diagnose whether a monthly or quarterly read is hiding a quality decline you are already paying for.
2. Design a plan and verify checkpoint for agent-driven changes, including where it sits in the workflow, who owns verification, and what it costs in developer time.
3. Estimate the true cost of agentic development for your organization, separating platform spend, which stabilizes, from tooling spend, which does not.
Enterprise Agents Without Chaos
Engineering teams can build impressive AI agent prototypes in a few days. The real challenge starts when that agent needs to call internal tools, retrieve company knowledge, respect user permissions, ask for approval, and leave an audit trail that security and compliance teams can trust. Enterprise-ready agentic capabilities are less about one clever prompt and more about platform architecture.
This session is for platform engineers who need to turn isolated agent experiments into reusable enterprise capabilities. We will walk through a practical reference architecture for an agent platform built around MCP servers, retrieval and context layers, memory, permissions, approval flows, monitoring, and evaluation. Instead of showing another single-agent demo, this session focuses on the platform decisions behind production use: what should be shared across teams, how to avoid duplicated implementations, how to prevent tool sprawl, and how to keep governance from becoming a delivery bottleneck.
Attendees will leave with a reference model for building agent capabilities that are reusable, governable, and safe to operate. They will also get a checklist for deciding what belongs in the platform, what belongs in each agent, and where human-in-the-loop is still required.
Your API Is Not Ready for Agents
APIs designed for human-driven applications often assume predictable workflows and deliberate actions. AI agents break those assumptions. They can select the wrong operation, submit ambiguous parameters, retry requests unexpectedly, or trigger irreversible side effects.
This session presents a practical approach to turning existing REST and event-driven APIs into reliable tools for AI agents. We will examine concrete failure modes in tool-calling applications and apply defensive patterns, such as explicit contracts, narrowly scoped permissions, idempotency, and human confirmation for sensitive operations.
A live demonstration will show the same agent interacting with two interfaces. Against a conventional API, it creates duplicate actions and exceeds its intended authority. Against a hardened tool interface, those failures are prevented, contained, or made reversible.
Rather than introducing another agent framework, the session focuses on reusable API and agentic architecture patterns. Attendees will leave with a checklist for evaluating existing endpoints and a practical guidance for designing agent-safe interfaces.
Engineering Greener Software
Software sustainability is often treated as a corporate reporting concern rather than an engineering responsibility. Yet architectural choices, inefficient workloads, unnecessary data movement, and over-provisioned infrastructure directly affect energy consumption, operating costs, and system performance.
This session presents a practical approach to reducing software’s environmental impact without compromising latency or reliability. We will examine how to measure energy consumption, identify carbon-intensive parts of a system, and evaluate environment-friendly improvements across application workloads.
Rather than promoting a particular platform or offering generic sustainability advice, we will examine the trade-offs between sustainability, performance, reliability, and cost. We will also discuss when sustainability and performance reinforce each other, when they conflict, and how to avoid optimizations that merely shift resource consumption to another part of the system.
Attendees will leave with a practical measurement approach and a decision framework for designing more sustainable real-world systems.
DSC DACH 2025
Common pitfalls in AI adoption
DATA BASH '25 Sessionize Event
move(data) 2025 Sessionize Event
DSC DACH 2024
AI Security for a safer tomorrow
IPS Open Day 2023
How can AI help you become a better developer?
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top