Call for Speakers

vLLM Toronto Meetup

in 40 days

vLLM Toronto Meetup

planned future dates

12 Nov 2026

location

Toronto, Canada


vLLM Toronto is a community meetup for engineers running LLM inference in production.

The first Toronto meetup was held at U of T's Schwartz Reisman Innovation Campus, hosted with NVIDIA, Red Hat, and the Vector Institute. Talks covered inference at scale, EAGLE speculative decoding, and Speculators. This is the second one, organized by the community.

Format: three or four talks in an evening, doors at 5:00 PM, networking after. Free to attend, open to everyone.

We want practitioner talks. Real deployments, real numbers, real failure modes. If you have run vLLM under load and learned something the docs do not cover, submit it.

Based in Toronto and the GTA. Speakers from anywhere who can get to the room are welcome.

This is a free, volunteer-organized community meetup. There is no entrance fee and no ticket of any kind. Attendance is open to anyone, and the call for speakers is open to anyone. Nobody is paid to organize it.

open, 11 months left
Call for Speakers
Call opens at 12:00 AM

07 Sep 2026

Call closes at 11:59 PM

05 Sep 2027

Call closes in Eastern Daylight Time (UTC-04:00) timezone.
Closing time in your timezone () is .

What we are looking for

Talks about vLLM itself — the engine, its internals, and what happens when you run it for real. Topics that fit:

  • vLLM internals: V1 engine, scheduler, model runner, how a model actually loads
  • KV cache and memory: prefix caching, KV connectors, offloading, LMCache, prefill/decode disaggregation
  • Performance: continuous batching, quantization (LLM Compressor, FP8/INT4, QAT), speculative decoding, torch.compile and CUDA graphs, custom kernels
  • Hardware backends: NVIDIA, AMD, Intel and Gaudi, TPU, and other accelerator plugins
  • Distributed serving: tensor, pipeline and expert parallelism, prefill/decode split, llm-d, vLLM Production Stack
  • Multimodal and vLLM-Omni: vision and audio serving, encoder cost, multimodal caching
  • Agentic workloads: tool calling, structured output, prefix reuse across turns
  • vLLM in RL and post-training: rollout engines, weight sync and hot reload
  • Model enablement: adding or fixing model support, day-0 upstream work
  • Benchmarking and observability: vllm bench, methodology, metrics, cost per token on live traffic
  • War stories: what broke in production, what it cost, what you changed

Formats (per speaker)

  • Full talk: 25 minutes plus 5 minutes Q&A
  • Lightning talk: 10 minutes, no Q&A

First-time speakers

Encouraged. If you have the material but not the stage time, say so in your submission and we will help you shape it.

Code of conduct

All speakers and attendees follow the vLLM community code of conduct.


Login with your preferred account


If you haven't logged in before, you'll be able to register.

Using social networks to login is faster and simpler, but if you prefer username/password account - use Classic Login.