Writing Mixture of Experts LLMs from Scratch in PyTorch
In this video, we discuss Mixture of Experts Transformers - the backbone behind popular LLMs like DeepSeek V3, Mixtral 8x22B, and more. You will learn concepts like Dense MOEs, Sparse MOEs, Top-K Routing, Noisy Routing, Expert Capacity, Switch Transformers, Auxilliary load balancing losses, and many more. Everything is presented visually to help conceptualize what is going on, and code snippets are provided to make it more concrete! Follow on Twitter: https://x.com/neural_avb To support this channel, you can buy me a coffee at: https://ko-fi.com/neuralavb Join the channel on Patreon to receive updates about the channel, and get access to bonus content used in all my videos. You will get the slides, notebooks, code snippets, word docs, and animations that went into producing this video. Here is the link: / neuralbreakdownwithavb Visit AI Agent Store Page: https://aiagentstore.ai/?ref=avishek #pytorch #transformers #deepseek Videos and playlists you would like: Attention to Transformers playlist: • Attention to Transformers from zero to her... Guide to fine-tuning open source LLMs: • Finetune LLMs to teach them ANYTHING with ... Generative Language Modeling from scratch: • From Attention to Generative Language Mode... References and additional links: Sparse Mixture of Experts paper: https://arxiv.org/abs/1701.06538 Mixtral of Experts: https://arxiv.org/abs/2401.04088 DeepSeek V2: https://arxiv.org/abs/2405.04434 DeepSeek V3: https://arxiv.org/abs/2412.19437 Switch Transformers / Expert Capacity: https://arxiv.org/abs/2101.03961 A Blog post: https://brunomaga.github.io/Mixture-o... A visual guide: https://newsletter.maartengrootendors... Survey paper: https://arxiv.org/pdf/2407.06204 Timestamps: 0:00 - Intro 1:52 - Mixture of Experts Intuition 4:53 - Transformers 101 9:20 - Dense MOEs 14:50 - Sparse MOEs 16:34 - Router Collapse and Top-K Routing 19:20 - Noisy TopK, Load Balancing 20:56 - Routing Analysis by Mixtral 22:30 - Auxilliary Losses & DeepSeek 24:05 - Expert Capacity 26:07 - 6 Points to Remember

Jürgen Klopp is the new national coach | Sportschau

Alexander Mercouris: NATO gerät bald in Panik und riskiert Krieg mit Russland

Building awesome Speech To Text Transformers from scratch - One line of Pytorch at a time!

Mixture of Experts (MoE), Visually Explained

How to finetune LLMs to THINK with Reinforcement Learning (GRPO from scratch!)

The Brain’s Learning Algorithm Isn’t Backpropagation
![Reverse Engineering Large Language Models | Build Your Own LLM Workshop #2 [Refreshed]](https://i.ytimg.com/vi/-bJs_F9COFE/hqdefault.jpg?sqp=-oaymwEjCNACELwBSFryq4qpAxUIARUAAAAAGAElAADIQj0AgKJDeAE=&rs=AOn4CLBCcxBBPnz2ZC5sQXYfVWqpAqZs6Q)
Reverse Engineering Large Language Models | Build Your Own LLM Workshop #2 [Refreshed]

Visualizing transformers and attention | Talk for TNG Big Tech Day '24

Steering LLM Behavior Without Fine-Tuning

Mixture of Experts: How LLMs get bigger without getting slower
![How Attention Got So Efficient [GQA/MLA/DSA]](https://i.ytimg.com/vi/Y-o545eYjXM/hqdefault.jpg?sqp=-oaymwEjCNACELwBSFryq4qpAxUIARUAAAAAGAElAADIQj0AgKJDeAE=&rs=AOn4CLBuOQf8Rw0rEDbSy5MucgJ2Vh6xGw)
How Attention Got So Efficient [GQA/MLA/DSA]

DeepSeek R1 Theory Overview | GRPO + RL + SFT
![How DeepSeek Rewrote the Transformer [MLA]](https://i.ytimg.com/vi/0VLAoVGf_74/hqdefault.jpg?sqp=-oaymwEjCNACELwBSFryq4qpAxUIARUAAAAAGAElAADIQj0AgKJDeAE=&rs=AOn4CLCSwSaI6q3w2_zizcjVK5wONqMqIQ)
How DeepSeek Rewrote the Transformer [MLA]

Transformers, the tech behind LLMs | Deep Learning Chapter 5

DeepSeekV4 - Manifold Constrained Hyper Connections (mHC) and the evolution of ResNets

A Visual Guide to Mixture of Experts (MoE) in LLMs

How Attention Mechanism Works in Transformer Architecture

Let me explain PyTorch in 7 Concepts

EASIEST Way to Train LLM Train w/ unsloth (2x faster with 70% less GPU memory required)

