But What Are Transformers?

Transformers is arguably the most influential neural network architecture in the last decade, powering the current boom of generative AI. In this video, we will review the basic ideas of the original encoder-decoder transformer architecture and understand how various design decisions are made. Enjoy! Slides download: https://www.dropbox.com/scl/fi/x7zkyd...

How Attention Got So Efficient [GQA/MLA/DSA]
▶︎

How Attention Got So Efficient [GQA/MLA/DSA]

A CPU Made of Atoms: IBM's Breakthrough 0.7nm Transistors
▶︎

A CPU Made of Atoms: IBM's Breakthrough 0.7nm Transistors

Rotary Position Embeddings (RoPE) Explained — The Rotation Trick Behind Long-Context LLMs
▶︎

Rotary Position Embeddings (RoPE) Explained — The Rotation Trick Behind Long-Context LLMs

Why are Transformers replacing CNNs?
▶︎

Why are Transformers replacing CNNs?

Transformers, the tech behind LLMs | Deep Learning Chapter 5
▶︎

Transformers, the tech behind LLMs | Deep Learning Chapter 5

Transformers: Attention Is Just Weighted Dot Products | The Math Behind AI
▶︎

Transformers: Attention Is Just Weighted Dot Products | The Math Behind AI

How Does the Transformer Encoder Actually Work? Complete Visual Breakdown
▶︎

How Does the Transformer Encoder Actually Work? Complete Visual Breakdown

The 60-Year Hunt for AI's Most Important Function
▶︎

The 60-Year Hunt for AI's Most Important Function

Richard Sutton – Father of RL thinks LLMs are a dead end
▶︎

Richard Sutton – Father of RL thinks LLMs are a dead end

Transformers Explained: The Discovery That Changed AI Forever
▶︎

Transformers Explained: The Discovery That Changed AI Forever

DeepSeek V4's Secret: 98% Less Memory
▶︎

DeepSeek V4's Secret: 98% Less Memory

The Algorithm That Made Modern AI Possible
▶︎

The Algorithm That Made Modern AI Possible

Visualizing transformers and attention | Talk for TNG Big Tech Day '24
▶︎

Visualizing transformers and attention | Talk for TNG Big Tech Day '24

Is Fine-Tuning Still Needed? LLMs, RAG, & LoRA
▶︎

Is Fine-Tuning Still Needed? LLMs, RAG, & LoRA

The Scariest Chart in Electrical Engineering
▶︎

The Scariest Chart in Electrical Engineering

Attention in transformers, step-by-step | Deep Learning Chapter 6
▶︎

Attention in transformers, step-by-step | Deep Learning Chapter 6

How Attention Mechanism Works in Transformer Architecture
▶︎

How Attention Mechanism Works in Transformer Architecture

It's about time we learn Transformers..
▶︎

It's about time we learn Transformers..

How might LLMs store facts | Deep Learning Chapter 7
▶︎

How might LLMs store facts | Deep Learning Chapter 7

Nobody Explained the Schrödinger Equation Like THIS!
▶︎

Nobody Explained the Schrödinger Equation Like THIS!

Ai companies are terrified
▶︎

Ai companies are terrified

Most AI has amnesia. Here's the fix
▶︎

Most AI has amnesia. Here's the fix

DeepSeek Gave LLMs a Real Memory (It's Not RAG)
▶︎

DeepSeek Gave LLMs a Real Memory (It's Not RAG)

Transformers: The best idea in AI | Andrej Karpathy and Lex Fridman
▶︎

Transformers: The best idea in AI | Andrej Karpathy and Lex Fridman

CHOSEN ONE!! YOUR IDENTITY REVEAL JUST SHOOK THE INTERNET... AND THEIR MINDS
▶︎

CHOSEN ONE!! YOUR IDENTITY REVEAL JUST SHOOK THE INTERNET... AND THEIR MINDS

Mixture of Experts (MoE), Visually Explained
▶︎

Mixture of Experts (MoE), Visually Explained

How Rotary Position Embedding Supercharges Modern LLMs [RoPE]
▶︎

How Rotary Position Embedding Supercharges Modern LLMs [RoPE]

We Tested $200 GPT-5.6 Sol on PhD Level Math
▶︎

We Tested $200 GPT-5.6 Sol on PhD Level Math

The Transformer Explained: A Complete Layer-by-Layer Visual Breakdown
▶︎

The Transformer Explained: A Complete Layer-by-Layer Visual Breakdown

Why AI Can Never Escape Turing's 1936 Proof
▶︎

Why AI Can Never Escape Turing's 1936 Proof