vLLM Toronto is a community meetup for engineers running LLM inference in production.
The first Toronto meetup was held at U of T's Schwartz Reisman Innovation Campus, hosted with NVIDIA, Red Hat, and the Vector Institute. Talks covered inference at scale, EAGLE speculative decoding, and Speculators. This is the second one, organized by the community.
Format: three or four talks in an evening, doors at 5:00 PM, networking after. Free to attend, open to everyone.
We want practitioner talks. Real deployments, real numbers, real failure modes. If you have run vLLM under load and learned something the docs do not cover, submit it.
Based in Toronto and the GTA. Speakers from anywhere who can get to the room are welcome.
This is a free, volunteer-organized community meetup. There is no entrance fee and no ticket of any kind. Attendance is open to anyone, and the call for speakers is open to anyone. Nobody is paid to organize it.
What we are looking for
Talks about vLLM itself — the engine, its internals, and what happens when you run it for real. Topics that fit:
Formats (per speaker)
First-time speakers
Encouraged. If you have the material but not the stage time, say so in your submission and we will help you shape it.
Code of conduct
All speakers and attendees follow the vLLM community code of conduct.
If you haven't logged in before, you'll be able to register.
Using social networks to login is faster and simpler, but if you prefer username/password account - use Classic Login.