Session
Inference Engineering with gRPC: Building Low-Latency AI Systems That Scale
Inference engineering is the biggest bottleneck in LLM-based AI systems today. Modern AI platform are increasingly relying on gRPC instead of REST, for communication between inference gateways, model servers, retrieval services, and orchestration layers because every millisecond matters.
This session introduces the engineering principles behind low-latency AI inference. It also highlights the tradeoffs between gRPC and REST as the protocol of choice for production LLM systems. Attendees will compare REST and gRPC for inference workloads, understand unary and streaming RPCs for real-time token generation, and learn how deadlines, cancellation, flow control, retries, and backpressure improve reliability under load. The session also explores how modern LLM serving stacks use gRPC to build scalable inference pipelines. Attendees will leave with practical guidance for designing faster, more resilient AI services using patterns they can immediately apply.
Smridhi Gupta
h Analyst @ Citibank | Building Trustworthy AI Systems | 7K+ Tech Community
Pune, India
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top