But What Are Transformers?
Transformers is arguably the most influential neural network architecture in the last decade, powering the current boom of generative AI. In this video, we will review the basic ideas of the original encoder-decoder transformer architecture and understand how various design decisions are made. Enjoy! Slides download: https://www.dropbox.com/scl/fi/x7zkyd...
![How Attention Got So Efficient [GQA/MLA/DSA]](https://i.ytimg.com/vi/Y-o545eYjXM/hqdefault.jpg?sqp=-oaymwEjCNACELwBSFryq4qpAxUIARUAAAAAGAElAADIQj0AgKJDeAE=&rs=AOn4CLBuOQf8Rw0rEDbSy5MucgJ2Vh6xGw)
▶︎
How Attention Got So Efficient [GQA/MLA/DSA]

▶︎
A CPU Made of Atoms: IBM's Breakthrough 0.7nm Transistors

▶︎
Rotary Position Embeddings (RoPE) Explained — The Rotation Trick Behind Long-Context LLMs

▶︎
Why are Transformers replacing CNNs?

▶︎
Transformers, the tech behind LLMs | Deep Learning Chapter 5

▶︎
Transformers: Attention Is Just Weighted Dot Products | The Math Behind AI

▶︎
How Does the Transformer Encoder Actually Work? Complete Visual Breakdown

▶︎
The 60-Year Hunt for AI's Most Important Function

▶︎
Richard Sutton – Father of RL thinks LLMs are a dead end

▶︎
Transformers Explained: The Discovery That Changed AI Forever

▶︎
DeepSeek V4's Secret: 98% Less Memory

▶︎
The Algorithm That Made Modern AI Possible

▶︎
Visualizing transformers and attention | Talk for TNG Big Tech Day '24

▶︎
Is Fine-Tuning Still Needed? LLMs, RAG, & LoRA

▶︎
The Scariest Chart in Electrical Engineering

▶︎
Attention in transformers, step-by-step | Deep Learning Chapter 6

▶︎
How Attention Mechanism Works in Transformer Architecture

▶︎
It's about time we learn Transformers..

▶︎
How might LLMs store facts | Deep Learning Chapter 7

▶︎
Nobody Explained the Schrödinger Equation Like THIS!

▶︎
Ai companies are terrified

▶︎
Most AI has amnesia. Here's the fix

▶︎
DeepSeek Gave LLMs a Real Memory (It's Not RAG)

▶︎
Transformers: The best idea in AI | Andrej Karpathy and Lex Fridman

▶︎
CHOSEN ONE!! YOUR IDENTITY REVEAL JUST SHOOK THE INTERNET... AND THEIR MINDS

▶︎
Mixture of Experts (MoE), Visually Explained
![How Rotary Position Embedding Supercharges Modern LLMs [RoPE]](https://i.ytimg.com/vi/SMBkImDWOyQ/hqdefault.jpg?sqp=-oaymwEjCNACELwBSFryq4qpAxUIARUAAAAAGAElAADIQj0AgKJDeAE=&rs=AOn4CLB6gWS_ZRO-UhithwlfNKgGNDFVNQ)
▶︎
How Rotary Position Embedding Supercharges Modern LLMs [RoPE]

▶︎
We Tested $200 GPT-5.6 Sol on PhD Level Math

▶︎
The Transformer Explained: A Complete Layer-by-Layer Visual Breakdown

▶︎
