He Co-Invented the Transformer. Now: Continuous Thought Machines [Llion Jones / Luke Darlow]

The Transformer architecture (which powers ChatGPT and nearly all modern AI) might be trapping the industry in a localized rut, preventing us from finding true intelligent reasoning, according to the person who co-invented it. Llion Jones and Luke Darlow, key figures at the research lab Sakana AI, join the show to make this provocative argument, and also introduce new research (CTM) which might lead the way forwards. We speak about "Inventor's Remorse" & The Trap of Success Despite being one of the original authors of the famous "Attention Is All You Need" paper that gave birth to the Transformer, Llion explains why he has largely stopped working on them. He argues that the industry is suffering from "success capture"—because Transformers work so well, everyone is focused on making small tweaks to the same architecture rather than discovering the next big leap. *SPONSOR MESSAGES START* — Build your ideas with AI Studio from Google - http://ai.studio/build — Tufa AI Labs is hiring ML Research Engineers https://tufalabs.ai/ — cyber•Fund https://cyber.fund/?utm_source=mlst is a founder-led investment firm accelerating the cybernetic economy Hiring a SF VC Principal: https://talent.cyber.fund/companies/c... Submit investment deck: https://cyber.fund/contact?utm_source... — *END* The "Spiral" Problem – Llion uses a striking visual analogy to explain what current AI is missing. If you ask a standard neural network to understand a spiral shape, it solves it by drawing tiny straight lines that just happen to look like a spiral. It "fakes" the shape without understanding the concept of spiraling. They argue that today's AI models are similar—they are incredible at mimicking intelligent answers without having an internal process of "thinking". Introducing the Continuous Thought Machine (CTM) Luke Darlow deep dives into their solution: a biology-inspired model that fundamentally changes how AI processes information. The Maze Analogy: Luke explains that standard AI tries to solve a maze by staring at the whole image and guessing the entire path instantly. Their new machine "walks" through the maze step-by-step. Thinking Time: This allows the AI to "ponder." If a problem is hard, the model can naturally spend more time thinking about it before answering, effectively allowing it to correct its own mistakes and backtrack—something current Language Models struggle to do genuinely. The pair discuss the culture of Sakana AI, which is modeled after the early days of Google Brain/DeepMind. Llion nostalgically recalls that the Transformer wasn't born from a corporate mandate, but from random people talking over lunch about interesting problems. https://sakana.ai/ https://x.com/YesThisIsLion https://x.com/LearningLukeD TRANSCRIPT: https://app.rescript.info/public/shar... TOC: 00:00:00 - Stepping Back from Transformers 00:00:43 - Introduction to Continuous Thought Machines (CTM) 00:01:09 - The Changing Atmosphere of AI Research 00:04:13 - Sakana’s Philosophy: Research Freedom 00:07:45 - The Local Minimum of Large Language Models 00:18:30 - Representation Problems: The Spiral Example 00:29:12 - Technical Deep Dive: CTM Architecture 00:36:00 - Adaptive Computation & Maze Solving 00:47:15 - Model Calibration & Uncertainty 01:00:43 - Sudoku Bench: Measuring True Reasoning REFS: Why Greatness Cannot be planned [Kenneth Stanley] https://www.amazon.co.uk/Why-Greatnes...    • Why Greatness Cannot Be Planned — Kenneth ...   The Hardware Lottery [Sara Hooker] https://arxiv.org/abs/2009.06489    • Sara Hooker - The Hardware Lottery, Sparsi...   Continuous Thought Machines [Luke Darlow et al / Sakana] https://arxiv.org/abs/2505.05522 https://sakana.ai/ctm/    • Continuous Thought Machine Deep Dive | Tem...   great walkthrough of algo by Yacine Mahdid LSTM: The Comeback Story? [Prof. Sepp Hochreiter]    • LSTM: The Comeback Story? [Prof. Sepp Hoch...   Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis [Kumar/Stanley] https://arxiv.org/pdf/2505.11581 Intelligent Matrix Exponentiation [Thomas Fischbacher] (Spiral reference) https://arxiv.org/abs/2008.03936 A Spline Theory of Deep Networks [Randall Balestriero] https://proceedings.mlr.press/v80/bal...    • Why LeCun Thinks Deep Learning Isn't Enoug...      • Neural Networks Are Elastic Origami! [Prof...   On the Biology of a Large Language Model [Anthropic, Jack Lindsey et al] https://transformer-circuits.pub/2025... The ARC Prize 2024 Winning Algorithm [Daniel Franzen and Jan Disselhoff] “The ARChitects”    • The ARC Prize 2024 Winning Algorithm [Dani...   Neural Turing Machine [Graves] https://arxiv.org/pdf/1410.5401 Adaptive Computation Time for Recurrent Neural Networks [Graves] https://arxiv.org/abs/1603.08983 Sudoko Bench [Sakana] https://pub.sakana.ai/sudoku/

Yann LeCun | Self-Supervised Learning, JEPA, World Models, and the future of AI
▶︎

Yann LeCun | Self-Supervised Learning, JEPA, World Models, and the future of AI

Google Researcher Shows Life "Emerges From Code" [Blaise Agüera y Arcas]
▶︎

Google Researcher Shows Life "Emerges From Code" [Blaise Agüera y Arcas]

Urgent Update- AI Sputnik Moment: Kimi K3 Released w/ Emad Mostaque | Ep. 272
▶︎

Urgent Update- AI Sputnik Moment: Kimi K3 Released w/ Emad Mostaque | Ep. 272

CHOSEN ONE!! YOUR IDENTITY REVEAL JUST SHOOK THE INTERNET... AND THEIR MINDS
▶︎

CHOSEN ONE!! YOUR IDENTITY REVEAL JUST SHOOK THE INTERNET... AND THEIR MINDS

David Krakauer on "Life, Intelligence, and the Many Worlds of Problem-Solving Matter"
▶︎

David Krakauer on "Life, Intelligence, and the Many Worlds of Problem-Solving Matter"

The Real Reason Huge AI Models Actually Work [Prof. Andrew Wilson]
▶︎

The Real Reason Huge AI Models Actually Work [Prof. Andrew Wilson]

This is why Deep Learning is really weird.
▶︎

This is why Deep Learning is really weird.

Creator of C++: Bell Labs, Negative Overhead Abstraction, Mistakes | Bjarne Stroustrup
▶︎

Creator of C++: Bell Labs, Negative Overhead Abstraction, Mistakes | Bjarne Stroustrup

Can Quantum Particles Communicate Faster Than Light? – Quantum Reality (3/3) with Jim Al-Khalili
▶︎

Can Quantum Particles Communicate Faster Than Light? – Quantum Reality (3/3) with Jim Al-Khalili

RL for Agents Workshop - Deep Dive on Training Agents with RL and Open Source
▶︎

RL for Agents Workshop - Deep Dive on Training Agents with RL and Open Source

Continuous Thought Machines and how to think about thought (Luke Darlow)
▶︎

Continuous Thought Machines and how to think about thought (Luke Darlow)

LLMs Don't Need More Parameters. They Need Loops.
▶︎

LLMs Don't Need More Parameters. They Need Loops.

François Chollet on OpenAI o-models and ARC
▶︎

François Chollet on OpenAI o-models and ARC

The Elegant Math Behind Machine Learning
▶︎

The Elegant Math Behind Machine Learning

The Thermodynamic AI Chip · Thomas Ahle
▶︎

The Thermodynamic AI Chip · Thomas Ahle

Keynote: After the AI Hype – What’s Real, and What’s Next - Richard Campbell - 2026
▶︎

Keynote: After the AI Hype – What’s Real, and What’s Next - Richard Campbell - 2026

Do LLMs Understand? AI Pioneer Yann LeCun Spars with DeepMind’s Adam Brown.
▶︎

Do LLMs Understand? AI Pioneer Yann LeCun Spars with DeepMind’s Adam Brown.

The Brain’s Learning Algorithm Isn’t Backpropagation
▶︎

The Brain’s Learning Algorithm Isn’t Backpropagation

By 2035, Most Grand Challenges Facing Humanity Will Be Solved w/ Dr. Alex Wissner-Gross
▶︎

By 2035, Most Grand Challenges Facing Humanity Will Be Solved w/ Dr. Alex Wissner-Gross

The Power of a Single Neuron and a Path to Simulating the Brain | Dr. Konrad Kording
▶︎

The Power of a Single Neuron and a Path to Simulating the Brain | Dr. Konrad Kording