Multi-Layer Perceptrons, Feed-Forward Networks | Build Your Own LLM Workshop #6

MLPs & Feedforward Networks: From a perceptrons, multiple inputs, multiple layers, and XOR to GPT-2-Style FFNs (ReLU², MatMul, PyTorch). Part of a Build your own LLM workshop. =========== LINKS Justin's twitter: https://x.com/JustinAngel Workshop overview: https://go.justinangel.ai/substack Deck: https://go.justinangel.ai/deck Google Drive: https://go.justinangel.ai/drive Code exercise: https://go.justinangel.ai/code-6 Excel exercise: https://go.justinangel.ai/excel-6 =========== CHAPTERS 00:00 Welcome and Goals 00:41 Weighted Sums Basics 01:56 Two Inputs Demo 03:55 Two Neurons and Weights 05:23 Adding Layers with ReLU 06:43 XOR Needs MLPs 09:05 MLP Terminology 10:10 Matmul in Sheets 13:46 Vectorized Multi Neuron Matmul 15:49 Bias Trick with Ones 17:06 XOR Table via Matmul 19:39 Growing Arrays with Matmul 20:37 Upscale Then Downscale 22:15 Matmul Dimension Rules 25:20 Excel Lessons Recap 26:47 Tensors In PyTorch 29:03 Why MLPs Expand 30:31 Build A PyTorch MLP 30:20 XOR In Code 32:54 Exercises And Parameters 35:56 GPT2 MLP Decisions 36:49 Playground Demos 38:53 Learn Linear Algebra 39:46 Wrap Up And Next =============== ABOUT THIS TALK Justin Angel reviews perceptrons and builds up to multilayer perceptrons (MLPs) and fully connected feedforward networks (FFNs) used in transformers, showing how multiple inputs sum as WX+B, how multiple neurons create more weights and biases, and how layers combine via a final output neuron with ReLU gating. He demonstrates why XOR requires multiple layers, then transitions from hand calculations to matrix multiplication (MMULT/matmul), including dimension rules, transposes, and a trick of folding biases into weights by appending ones to inputs. He connects this to LLM practice of expanding and contracting dimensions (e.g., 4x upsample then downsample) as a scratchpad, cites Geva (2021), and implements MLPs in PyTorch (including XOR) plus an exercise to build a GPT-2-style FFN: 768 → 3072 → 768 with ReLU squared and typically no biases.