Session

Attention Attention Everywhere

Everyone knows the line "attention is all you need". Far fewer could say what attention actually computes.

This talk starts from a single sentence with an ambiguous pronoun and works forward until the attention formula appears on its own. Not presented and explained, but derived, one step at a time, with real numbers on screen throughout. Why comparing two embeddings directly fails. Why that failure forces every token to be split into three separate projections. What each projection does and what actually moves between tokens.

Then it stops explaining and starts checking. The same math is run against the real weights of an open model, pulled out layer by layer, head by head, and recomputed by hand. The first measurement comes back wrong, and why it comes back wrong turns out to be more interesting than the mechanism itself.

Dev J. Shah

SWE @TribalScale, GenAI Evangelist (Blogger, Speaker) || 4x Multi-Cloud Certified || Software Engineering, AI Engineering || Linux, Cloud, DevOps

Toronto, Canada

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top