Is Kimi K3 Really That Good?! (Don't Just Believe The Hype)
Kimi K3 is the most powerful open weight model ever released, and the benchmarks make it look like it's beating Opus 4.8 and almost matching the very best closed models. After spending millions of tokens putting it up against Opus 4.8 and Kimi K2.7 on real engineering tasks, I'm not buying it, and I don't think you should either. Kimi K3 is genuinely impressive, and sometimes the raw output is even better than Opus. But open weight models have reliability failure modes that the public benchmarks don't seem to account for. In this video I show you exactly where Kimi K3 breaks down, why it's still useful as the workhorse in a mixed-model workflow, and how to build your own benchmarks that actually reflect real world coding instead of a leaderboard. Everything is open source so you can run these tests yourself, even with other models! ~~~~~~~~~~~~~~~~~~~~~~~~~~ QA.tech, the AI testing tool with autonomous QA agents that test your app like a real user: http://qa.tech/cole Past the Bottleneck - free guide to releasing quality products in the AI-SDLC, for teams that ship fast and want to stay safe: http://qa.tech/cole-book ~~~~~~~~~~~~~~~~~~~~~~~~~~ Join my free live workshop on July 29th, where I show you how to become an AI native engineering organization - building a reliable standard for how teams use AI coding agents: https://dynamous.ai/ai-native-enginee... Benchmark repo (workflows, rubric, and every exact prompt): https://github.com/coleam00/kimi-k3-r... Archon (the open source harness builder I used): https://github.com/coleam00/archon ~~~~~~~~~~~~~~~~~~~~~~~~~~ 0:00 The Kimi K3 Hype vs Reality 1:41 The Benchmark Suite I Built 3:43 Why Open Weight Models Are Worth It 4:47 Mixing Models for a Cheaper Workflow 5:50 Benchmark 1: Real Engineering Tasks 8:16 Sponsor: QA.tech 10:19 How the Scoring Works (7 Dimensions) 10:56 Simple Tasks: K3 Basically Ties Opus 12:28 Complex Tasks: Opus Pulls Ahead 14:17 Why Public Benchmarks Can't Be Trusted 15:33 The Trap Tasks Explained 17:00 The Results: 8% vs 36% Failure Rate 18:58 Why Opus Thinks for Itself 20:12 The Takeaway: Plan Big, Build Cheap ~~~~~~~~~~~~~~~~~~~~~~~~~~ Join me as I push the limits of what is possible with AI. I'll be uploading videos weekly - at least every Wednesday at 7:00 PM CDT!

OPUS 5 CLICK NOW

Elon Musk speaks to 'The Economist': Editor-in-chief Zanny Minton Beddoes on the key takeaways

Watch an AI Team Mate Explore a GitHub Repository Using MCP | TruGen AI

The Best AI Coding Setup Isn't the Most Autonomous One (Here's Why)

Claude Opus 5 review: great at coding (but I hate talking to it)

Opus 5 Is Here And Its BETTER THAN FABLE?!

The 3 elements of stupidity, according to philosophy | Jonny Thomson: Full Interview

Vibe Coding With Claude Opus 5

Blacking Out a Word Doc Doesn't Actually Delete Anything. Do This Instead.

A380 Hydraulic Failure: KAL81 Needed a Manual Gear Extension

Harness Engineering: What Separates Top Agentic Engineers Right Now

It Took Me 6 Years to Make This

Straight Talk # 215

Claude Opus 5 is Going to Save You Money

First words as head coach: Klopp issues a clear warning!

Google Just Dropped a Masterclass on Agentic Engineering (It's SO Good)

China Just Built an AI Chip Out of Light (To Bypass Silicon Limits)

Energiekrise - Update und Resilienz durch Speicher

Elon Musk on the Future of Tesla FSD - V15 and More

