Session
Knowledge Distillation: How LLMs train SLMs
Gemma was trained by Gemini. Llama 4's smaller models were trained by their two-trillion-parameter sibling. DeepSeek does something in the same family. The technique is called knowledge distillation.
This session traces distillation's near-20-year history — from compressing thousand-model ensembles onto PDAs in 2006, to Hinton's "dark knowledge" and the teacher–student framing, to the temperature knob you use every day. We'll then look at how Google, Meta, and DeepSeek distill their models today, and why the implementation details — proper distillation vs. behavioral cloning, who owns the teacher, sequential vs. co-training — quietly make or break a model.
Attendees will leave understanding what distillation actually transfers between models, why "soft labels" carry more than answers, and how to tell genuine distillation from mere imitation.
Rama Krishna Raju Samantapudi
Sr. Staff AI/ML Architect at ServiceNow
Austin, Texas, United States
Links
Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.
Jump to top