Call for Speakers

vLLM Toronto Meetup

in 2 months

vLLM Toronto Meetup

planned future dates

12 Nov 2026

location

Toronto, Canada


vLLM Toronto is a community meetup for engineers running LLM inference in production.

The first Toronto meetup was held at U of T's Schwartz Reisman Innovation Campus, hosted with NVIDIA, Red Hat, and the Vector Institute. Talks covered inference at scale, EAGLE speculative decoding, and Speculators. This is the second one, organized by the community.

Format: three or four talks in an evening, doors at 5:00 PM, networking after. Free to attend, open to everyone.

We want practitioner talks. Real deployments, real numbers, real failure modes. If you have run vLLM under load and learned something the docs do not cover, submit it.

Based in Toronto and the GTA. Speakers from anywhere who can get to the room are welcome.

This is a free, volunteer-organized community meetup. There is no entrance fee and no ticket of any kind. Attendance is open to anyone, and the call for speakers is open to anyone. Nobody is paid to organize it.

open, 12 months left
Call for Speakers
Call opens at 12:00 AM

07 Sep 2026

Call closes at 11:59 PM

05 Sep 2027

Call closes in Eastern Daylight Time (UTC-04:00) timezone.
Closing time in your timezone () is .

What we are looking for

Talks from people who operate inference systems. Topics that fit:

  • vLLM in production: scaling, scheduling, autoscaling, multi-tenancy
  • KV cache: offloading, disaggregation, prefix caching, remote pools
  • Serving performance: batching, quantization, speculative decoding, kernel work
  • Distributed serving: prefill/decode split, P/D routing
  • Kubernetes and vLLM: operators, Helm, KEDA, GPU sharing and scheduling
  • Model routing, gateways, and routing policy
  • Evaluation, observability, and cost work on live traffic
  • War stories: what broke, what it cost, what you changed

Formats

  • Full talk: 25 minutes plus 5 minutes Q&A
  • Lightning talk: 10 minutes, no Q&A

Level

Intermediate to advanced. Assume the room has deployed something and knows what a KV cache is. Benchmarks with stated methodology beat a single headline number.

What does not fit

Product pitches. Sponsors are welcome and credited, but the stage is for technical content. If your talk cannot be given without your company's product, it is not a fit.

First-time speakers

Encouraged. If you have the material but not the stage time, say so in your submission and we will help you shape it.

Code of conduct

All speakers and attendees follow the vLLM community code of conduct.


Login with your preferred account


If you haven't logged in before, you'll be able to register.

Using social networks to login is faster and simpler, but if you prefer username/password account - use Classic Login.