Attention, KV Cache, MQA & GQA — A Visual Guide

A visual deep-dive into how attention works in modern LLMs — from embeddings and Q, K, V projections to KV caching, Multi-Query Attention, and Grouped-Query Attention. No prior deep learning knowledge required. Some diagrams in this video were inspired by Under The Hood — go check out their channel! 📌 Chapters: 0:00 Introduction 1:34 Embeddings 3:10 Attention Formula 4:10 Queries, Keys & Values 7:43 Scaled Dot-Product Attention 9:40 Masked Self-Attention 11:50 Multi-Head Attention 13:20 KV Caching 15:10 Multi-Query Attention (MQA) 16:05 Grouped-Query Attention (GQA) 16:30 Summary Music by Vincent Rubinetti Download the music on Bandcamp: https://vincerubinetti.bandcamp.com Stream the music on Spotify: https://open.spotify.com/artist/2SRhEEt2tl...