I Visualised Attention in Transformers
To try everything Brilliant has to offer—free—for a full 30 days, visit https://brilliant.org/GalLahat/ . You’ll also get 20% off an annual premium subscription. Voice type with Peach Beta 🍑: https://peach-voice.com This video was sponsored by Brilliant The music is created by my partner (AI) and me, feel free to use it commercially for your own projects but make sure to credit this video when you do: https://drive.google.com/drive/folder...

▶︎
Attention in transformers, step-by-step | Deep Learning Chapter 6

▶︎
Visualizing transformers and attention | Talk for TNG Big Tech Day '24

▶︎
The math behind Attention: Keys, Queries, and Values matrices
![Yann LeCun's $1B Bet Against LLMs [Part 1]](https://i.ytimg.com/vi/kYkIdXwW2AE/hqdefault.jpg?sqp=-oaymwEjCNACELwBSFryq4qpAxUIARUAAAAAGAElAADIQj0AgKJDeAE=&rs=AOn4CLDbV4izF3i-wxevCVIn7FJjoy1vlA)
▶︎
Yann LeCun's $1B Bet Against LLMs [Part 1]

▶︎
Transformers: Attention Is Just Weighted Dot Products | The Math Behind AI
![How DeepSeek Rewrote the Transformer [MLA]](https://i.ytimg.com/vi/0VLAoVGf_74/hqdefault.jpg?sqp=-oaymwEjCNACELwBSFryq4qpAxUIARUAAAAAGAElAADIQj0AgKJDeAE=&rs=AOn4CLCSwSaI6q3w2_zizcjVK5wONqMqIQ)
▶︎
How DeepSeek Rewrote the Transformer [MLA]

▶︎
Why AI Has Failed to Take Your Job Since 1976

▶︎
How does AI actually work? Transformers explained

▶︎
Transformers, the tech behind LLMs | Deep Learning Chapter 5

▶︎
LLMs Don't Need More Parameters. They Need Loops.

▶︎
Recursive Self-Improvement

▶︎
Hermes Agent Fundamentals In 29 Minutes

▶︎
Query, Key and Value Matrix for Attention Mechanisms in Large Language Models

▶︎
MIT Explains the 12 Possible Endings for AI

▶︎
Keys, Queries, and Values: The celestial mechanics of attention
![How Attention Got So Efficient [GQA/MLA/DSA]](https://i.ytimg.com/vi/Y-o545eYjXM/hqdefault.jpg?sqp=-oaymwEjCNACELwBSFryq4qpAxUIARUAAAAAGAElAADIQj0AgKJDeAE=&rs=AOn4CLBuOQf8Rw0rEDbSy5MucgJ2Vh6xGw)
▶︎
How Attention Got So Efficient [GQA/MLA/DSA]

▶︎
Squares are 83.2% circle

▶︎
I Forced AI To Learn 4D Movement

▶︎
Attention is all you need (Transformer) - Model explanation (including math), Inference and Training

▶︎
