Session
AI at the Gateway: Scaling, Securing, and Orchestrating Agentic AI
Integrating LLMs into production is easy for a demo, but scaling them for thousands of concurrent users reveals a harsh reality: they are high-latency, expensive, and fundamentally non-deterministic. Traditional API Gateways weren't built for 60-second request times or "hallucinating" status codes.
In this session, we move beyond basic API calls to explore the AI Gateway Pattern. We will dive into how to build a resilient orchestration layer that treats AI models as unreliable downstream dependencies rather than standard microservices. You will see how to implement "semantic caching" to save costs, handle token-based rate limiting, and use circuit breakers to failover between disparate models (e.g., from GPT-4 to a local Llama instance).
We will walk through the transition from synchronous "Request/Response" to Event-Driven AI orchestration, ensuring your architecture remains responsive even when the model is slow. This is a session about the "plumbing" of AI—security, observability, and stability—designed for engineers who need to move AI from a playground to a production-grade system.
Hugo Guerrero
Building the Infrastructure for the Agentic Era | AI, MCP, Kubernetes & Cloud Native | Speaker on AI, APIs & AX/DX
Boston, Massachusetts, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top