Alibaba Claimed #2 In The World. I Tested It. They Lied.
Alibaba dropped Qwen 3.8 Max in preview, a 2.4 TRILLION parameter model, and claimed it's second only to Claude Fable 5. That is an enormous claim. So I ran it through all nine tests on my coding benchmark and scored every single one live on camera, no edits, no idea what the results were before I hit record. It did NOT land where they said it would. The coding tests were flawless. The skepticism audit was genuinely impressive, it caught every fake problem when Kimi K3 and GLM 5.2 both failed. But the Halo clone had broken health and shield gauges, the XCOM clone put enemies inside walls, and the 3D printable Tardis got the proportions wrong when every other frontier model got it right. Final weighted score: 93. Fourth place. Behind Kimi K3, GPT 5.6 Soul, and Claude Fable 5. Also, a tip if you want to try it: the personal coding plan link is basically unfindable on Alibaba Cloud. I had to ask their own AI agent to hand me the URL. That actually worked. CHAPTERS 0:00 Alibaba's huge claim about Qwen 3.8 Max 0:47 It's coding-plan only (and the signup link is hidden) 1:27 The nine tests on my benchmark, explained 2:58 Test 1: Halo Arena clone 4:30 Test 2: XCOM tactical grid 6:10 Test 3: 3D printable object (the hardest test) 7:24 Test 4: Sales page and animated SVGs 8:47 Coding, reasoning traps, and the skepticism audit 9:29 Grid Forge interactive spreadsheet 10:17 The final score and the leaderboard 12:01 Side by side vs Kimi K3 and Fable 5 15:34 Pricing, and should you actually use it? LINKS My site: https://mattjohnston.io I test every new model on this same benchmark the day it drops, so subscribe if you want the real numbers instead of the launch-day marketing.

KIMI K3 is Ridiculous.

The Alibaba AI Incident Should Terrify Us - Tristan Harris

Qwen3.8 MAX Preview Is HERE – Is THIS the BEST Open Model Yet?

Ternary Bonsai 27B benchmarked and tested vs Qwen 27B - 16GB Local LLM setup

The AI Backlash Could Actually Be Lethal

Stanislav Krapivnik: Russlands Wut kocht über – Steht ein EU-Russland-Krieg bevor?

UK Users Are Switching to Tor Browser — Here's What It Actually Does

The World's Smartest People Are Sprinting for the Same Exit

The Non-NVIDIA AI Card Everyone’s Ignoring

LinusTechTips Is Losing Thousands of Fans. Why?

China's K3 Model Reveals the Problem With Open Weights

I Gave Kimi K3, GPT-5.6 and Fable 5 the Same 9 Game Tests

The truth about Bun’s Rust rewrite

Qwen 4.0 LEAKED, DeepSeek V4 GA Launching Monday? & GLM 5.3 Soon! HUGE AI NEWS

Helios Is AMD’s First AI System To Rival Nvidia Vera Rubin — We Got An Exclusive, First Look

I Tested Qwen 3.8 So You Don't Have To…

Jiang: 90% of Humanity Could Be Gone in 50 Years. What Could Cause It?

Open Source AI Is Getting Too Big to Run

Kimi K3 Delivers Frontier AI at 1% of the Cost: AI Sputnik Moment w/ Emad Mostaque | Ep. 272

