The most complex model we actually understand
New AI Book! https://www.welchlabs.com/resources/a... Get a free ebook version today when you order a copy from our January 2026 print run! You’ll receive a discount code for 100% off the ebook in your purchase confirmation email. ebook: https://www.welchlabs.com/resources/t... Patreon: / welchlabs Sections 0:00 - Intro 2:39 - Modular Addition 3:54 - The Model’s Perspective 6:52 - An Accidental Discovery at OpenAI 7:56 - It Groks! 8:49 - Some Clues 13:17 - New Welch Labs Book! 13:57 - Deeper into the model 15:13 - Linear Probes 16:59 - Clocks perform modular addition 19:17 - How do x and y interact exactly? 23:19 - It learns a trig identity?! 26:38 - Putting the pieces together & excluded loss 30:02 - Anthropic finds 6D manifolds 32:24 - Final thoughts 34:02 - Welch Labs update Special thanks to Neel Nanda for discussing his work and Mech Interp with me, if you want to learn more about Mech Interp, check out Neel’s getting started post here: https://neelnanda.io/getting-started Thanks to Emmanuel Ameisen and Wes Gurnee for discussing their work on Claude Haiku. Their paper is incredibly in depth and interesting: https://transformer-circuits.pub/2025... Really nice deeper breakdown on polynomial double descent from viewer Avaneesh Narla: https://www.avaneeshnarla.com/blog/do... Interesting Grokking Analysis from viewer Anthony Z: https://huggingface.co/spaces/zborals... OpenAI team’s grokking paper: https://arxiv.org/pdf/2201.02177. I wasn’t able to reach to team for comment on the origin story, but it is told here: • Girl Geek X OpenAI Lightning Talks Nanda et al. https://arxiv.org/pdf/2301.05217v1 More on Grokking: https://www.quantamagazine.org/how-do... Code based on excellent these notebooks from Neel Nanda and collaborators: https://colab.research.google.com/dri... Andrej Karapathy on “Summoning Ghosts”: https://karpathy.bearblog.dev/animals... Code: https://github.com/stephencwelch/mani... Technical Notes It’s very natural for the attention layer to take the sum of it’s inputs (e.g. cos(kx)+cos(ky)), however we also find strong product terms. There’s a couple ways the network can compute products like cos(kx)cos(ky). One option is to approximate the product using ReLU activation functions (see Nanda’s notebooks for more). It’s also feasible for the attention block to do this, I found evidence of this is my own exploration In the first 2D fourier decomposition, we’re leaving out one component, specifically a “negative frequency component” → 0.26 * np.cos(2*np.pi*((4*i)/113)) * np.cos(2*np.pi*((109*j)/113)). We left this out to avoid digging into a discussion of negative/aliased frequencies, and having this 4th component doesn’t add to our intuition about what the network is doing here. at 28:00 we’re not showing removing the 8pi/113 frequency from the model’s final output surface. Patrons Juan Benet, Ross Hanson, Yan Babitski, AJ Englehardt, Alvin Khaled, Eduardo Barraza, Hitoshi Yamauchi, Jaewon Jung, Mrgoodlight, Shinichi Hayashi, Sid Sarasvati, Dominic Beaumont, Shannon Prater, Ubiquity Ventures, Matias Forti, Brian Henry, Tim Palade, Petar Vecutin, Nicolas baumann, Jason Singh, Robert Riley, vornska, Barry Silverman, Jake Ehrlich, Mitch Jacobs, Lauren Steely, Jeff Eastman, Rodolfo Ibarra, Clark Barrus, Rob Napier, Andrew White, Richard B Johnston, abhiteja mandava, Burt Humburg, Kevin Mitchell, Daniel Sanchez, Ferdie Wang, Tripp Hill, Richard Harbaugh Jr, Prasad Raje, Kalle Aaltonen, Midori Switch Hound, Zach Wilson, Chris Seltzer, Ven Popov, Hunter Nelson, Amit Bueno, Scott Olsen, Johan Rimez, Shehryar Saroya, Tyler Christensen, Beckett Madden-Woods, Darrell Thomas, Javier Soto, U007D, Caleb Begly, Rick Rubenstein, Brent Hunsaker, Dan Patterson, Tchsurvives, Alex Adai, Walter Reade, Zyansheep, Walter Reade, Duncan Stannett, Reginald Carey, Jean-Manuel Izaret, dh71633, Adrian Rodriguez, Dimitar Stojanovski, Michael Harder, Peter Maldonado, Emily Pesce, David Johnston, Insang Song, FaeTheWolf, Stephen Taylor, KittenKaboodle, EMatter, PATRICKMCCORMACK, John Beahan, Cameron, Cole Jones, Garrett Thornburg, Jeroen W, Rohit Sharma, GlennB, Emmanuel Cortes, Katie Quinn, Karina C, Cakra WW, Mike Ton, Eric Gometz, MacCallister Higgins, Niko Drossos, David Eraso, Tom Zehle, Steve, Brian Lineburg, rjbl, Michael Loh, Perry Vais, Bengal0, Farhad Manjoo, Sara Chipps, Ellis Driscoll, William Taysom, Will Harmon, CK, Abdullah, Peter Cho, Leo Nikora, Griffin Smith, Ash Katnoria, Alex, Markus Hays Nielsen, Catherine H., Vi, David Dobáš, Peter Wang, Sina Sohangir, Danny Thomas, Julian Francis, Hans Adler, Jiayu Peng Created by: Stephen Welch, Sam Baskin, and Pranav Gundu CFAQJOTYQHT7JYIT

The REAL reason the US can’t beat Iran

Finally: Grokking Solved - It's Not What You Think
![Yann LeCun's $1B Bet Against LLMs [Part 1]](https://i.ytimg.com/vi/kYkIdXwW2AE/hqdefault.jpg?sqp=-oaymwEjCNACELwBSFryq4qpAxUIARUAAAAAGAElAADIQj0AgKJDeAE=&rs=AOn4CLDbV4izF3i-wxevCVIn7FJjoy1vlA)
Yann LeCun's $1B Bet Against LLMs [Part 1]

Nobody Explained the Schrödinger Equation Like THIS!

Neural Networks Explained — What They Are & How They Actually Think

The Scariest Chart in Electrical Engineering
![Why Deep Learning Works Unreasonably Well [How Models Learn Part 3]](https://i.ytimg.com/vi/qx7hirqgfuU/hqdefault.jpg?sqp=-oaymwEjCNACELwBSFryq4qpAxUIARUAAAAAGAElAADIQj0AgKJDeAE=&rs=AOn4CLBuo8VfTixNDA_9nS8hYRCQvGpFtg)
Why Deep Learning Works Unreasonably Well [How Models Learn Part 3]

The Most Important Conversation in AI Right Now

I Built an LLM From Scratch

The Most Controversial Idea In Physics

But what is quantum computing? (Grover's Algorithm)

Recursive Self-Improvement

Why AI Tokens are so Expensive - Computerphile

LLMs Don't Need More Parameters. They Need Loops.
![Inside the World's Smartest Robot Brain [VLA]](https://i.ytimg.com/vi/2mrGMMmrVNE/hqdefault.jpg?sqp=-oaymwEjCNACELwBSFryq4qpAxUIARUAAAAAGAElAADIQj0AgKJDeAE=&rs=AOn4CLBsKsDGBMWqIscKPoqdk4iZtfZeeQ)
Inside the World's Smartest Robot Brain [VLA]

Transformers, the tech behind LLMs | Deep Learning Chapter 5

Can humans make AI any better?

Andrej Karpathy: From Vibe Coding to Agentic Engineering w/ Stephanie Zhan

But how do AI images and videos actually work? | Guest video by Welch Labs

