What We Didn't Cover | Build Your Own LLM Workshop #23

Scaling, GPU Coding, Flash Attention, KV Caching, Inference, and Safety. Part of a Build your own LLM workshop. =========== LINKS Justin's twitter: https://x.com/JustinAngel Workshop overview: https://go.justinangel.ai/substack Deck: https://go.justinangel.ai/deck Google Drive: https://go.justinangel.ai/drive =========== CHAPTERS 00:00 Final Section Intro 00:32 Scaling Up Overview 01:21 Parallelism and MoE 02:31 Flash Attention Basics 03:27 Training Scale Techniques 04:19 Memory and Sharding Tricks 05:33 Mixed Precision and Distillation 06:20 Inference Scaling Toolbox 07:18 KV Cache and Decoding Speedups 09:30 AI Safety Landscape 11:47 Practical Training Decisions 13:41 Wrap Up and Next Steps =============== ABOUT THIS TALK In the final mini chapter of the “Build Your Own LLM” workshop, the speaker reviews key topics not covered in depth, focusing on what learners should study next to scale training and inference. They survey techniques such as tensor, expert, and pipeline parallelism; mixture-of-experts; flash attention; gradient checkpointing; sharding approaches like FSDP/ZeRO; sequence and context parallelism; mixed precision; and distillation. For inference, they highlight quantization, KV caching, speculative decoding, multi-node inference, and prefill/decode disaggregation, recommending specific books and resources. The episode also outlines major AI safety areas including alignment, mechanistic interpretability, security/red teaming, unlearning/editing, activation steering, and evaluations, and notes practical training choices like TF32/BF16 and flash attention versions. The speaker closes by summarizing skills gained and emphasizing GPU proficiency for hireability.