What We Didn't Cover | Build Your Own LLM Workshop #23
Scaling, GPU Coding, Flash Attention, KV Caching, Inference, and Safety. Part of a Build your own LLM workshop. =========== LINKS Justin's twitter: https://x.com/JustinAngel Workshop overview: https://go.justinangel.ai/substack Deck: https://go.justinangel.ai/deck Google Drive: https://go.justinangel.ai/drive =========== CHAPTERS 00:00 Final Section Intro 00:32 Scaling Up Overview 01:21 Parallelism and MoE 02:31 Flash Attention Basics 03:27 Training Scale Techniques 04:19 Memory and Sharding Tricks 05:33 Mixed Precision and Distillation 06:20 Inference Scaling Toolbox 07:18 KV Cache and Decoding Speedups 09:30 AI Safety Landscape 11:47 Practical Training Decisions 13:41 Wrap Up and Next Steps =============== ABOUT THIS TALK In the final mini chapter of the “Build Your Own LLM” workshop, the speaker reviews key topics not covered in depth, focusing on what learners should study next to scale training and inference. They survey techniques such as tensor, expert, and pipeline parallelism; mixture-of-experts; flash attention; gradient checkpointing; sharding approaches like FSDP/ZeRO; sequence and context parallelism; mixed precision; and distillation. For inference, they highlight quantization, KV caching, speculative decoding, multi-node inference, and prefill/decode disaggregation, recommending specific books and resources. The episode also outlines major AI safety areas including alignment, mechanistic interpretability, security/red teaming, unlearning/editing, activation steering, and evaluations, and notes practical training choices like TF32/BF16 and flash attention versions. The speaker closes by summarizing skills gained and emphasizing GPU proficiency for hireability.

Using Large Language Models | Build Your Own LLM Workshop #1

Multi-Layer Perceptrons, Feed-Forward Networks | Build Your Own LLM Workshop #6

How AI Cracked an "Unsolvable" 87-Year Math Puzzle in Hours

I Built an LLM From Scratch

Kimi K3 explained in 13min..

Context Management masterclass | Technical Deep Dive Managing Context Agentic Systems

TypeScript 7 Is Here (And It's 10× Faster)

Is This Wish Meant to Be Fulfilled? 🧚🤲 Detailed Pick a Card Tarot Reading ✫・

Chosen One, This Is Why God Kept You Single All This Time

How Nvidia GPUs Compare To Google’s And Amazon’s AI Chips

No Boss, No Money: The Raw Reality of China’s Gen-Z Freelancers

Selling Dr*gs at my Job in CHICAGO in GTA 5 RP

How is hardware reshaping LLM design?
![Activation Functions: ReLU, GELU | Build Your Own LLM Workshop #4 [Refreshed]](https://i.ytimg.com/vi/ozvUxyDJvl4/hqdefault.jpg?sqp=-oaymwEjCNACELwBSFryq4qpAxUIARUAAAAAGAElAADIQj0AgKJDeAE=&rs=AOn4CLAGxA1zsnQyWMmbnO_J53k_-zwS8A)
Activation Functions: ReLU, GELU | Build Your Own LLM Workshop #4 [Refreshed]

DO NOT BUY: LG’s Spyware TVs, Monitors, and Wiretapping Concerns

URGENT UPDATE - Iran War Expert: A Mass Casualty Attack Is Coming! | Robert Pape
![Reverse Engineering Large Language Models | Build Your Own LLM Workshop #2 [Refreshed]](https://i.ytimg.com/vi/-bJs_F9COFE/hqdefault.jpg?sqp=-oaymwEjCNACELwBSFryq4qpAxUIARUAAAAAGAElAADIQj0AgKJDeAE=&rs=AOn4CLBCcxBBPnz2ZC5sQXYfVWqpAqZs6Q)
Reverse Engineering Large Language Models | Build Your Own LLM Workshop #2 [Refreshed]

How AI agents & Claude skills work (Clearly Explained)

