Session

Self-Hosting Agents: What Changes When Models Become Infrastructure

Running open models locally changes how engineers think about agent systems. Once models move from hosted APIs into self-managed infrastructure, problems that were previously hidden behind providers become operational realities: latency, orchestration, throughput, observability, GPU memory, tool-call reliability, evaluation variance, and cost control.

This talk shares lessons learned building and operating self-hosted agent systems with llama.cpp, vLLM, quantized models, local coding agents, and production-like homelab infrastructure. Using examples from Pedro CLI, multi-model serving, evaluation workflows, and distributed inference experiments, we will look at what changes when the model is no longer an API call but part of the platform.

Rather than focusing on benchmarks or hype, this session focuses on the operational work required to make open models useful in agent workflows: routing, retries, context management, observability, quantization tradeoffs, and failure recovery.

Attendees will leave with a practical model for deciding when self-hosting makes sense, what infrastructure problems to expect, and how to design open agent systems that are reliable enough to operate.

Miriah Peterson

Building context infrastructure for reliable AI systems | Data Engineering | Agentic Data Layers | A DomesticatingAi Podcast co-host

Salt Lake City, Utah, United States

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top