Cost-Efficient AI: Why Inkling's 975B MoE Changes Everything. SWE-1.7, Condensed Chain-of-Thought.

What if the secret to the next generation of AI isn't just building bigger data centers, but fundamentally changing how models think—and how we train them across the globe? We are unpacking a massive drop from the AI frontier: the newly released Inkling model card from Thinking Machines Lab. At 975 billion parameters, Inkling isn't just another massive Mixture-of-Experts. It is a natively multimodal engine built from the ground up for a new era of cost-efficient intelligence. We’re going to look at how this beast acts as the foundation for SWE-1.7, pushing the absolute limits of agentic software engineering and long-horizon tasks. But the real story here is the engineering underneath. We are breaking down the wild logistics of their fault-tolerant, distributed reinforcement learning pipeline spanning three continents. We’ll explore the math behind their gradient norm stabilization, the use of the Muon optimizer to keep the training run alive, and a fascinating new capability called condensed chain-of-thought—where the AI literally self-compacts its own reasoning to slash compute costs. Whether you're building autonomous agents or just trying to keep up with the bleeding edge of AI architecture, strap in. Let's dive into the Inkling framework.