Alibaba Claimed #2 In The World. I Tested It. They Lied.

Alibaba dropped Qwen 3.8 Max in preview, a 2.4 TRILLION parameter model, and claimed it's second only to Claude Fable 5. That is an enormous claim. So I ran it through all nine tests on my coding benchmark and scored every single one live on camera, no edits, no idea what the results were before I hit record. It did NOT land where they said it would. The coding tests were flawless. The skepticism audit was genuinely impressive, it caught every fake problem when Kimi K3 and GLM 5.2 both failed. But the Halo clone had broken health and shield gauges, the XCOM clone put enemies inside walls, and the 3D printable Tardis got the proportions wrong when every other frontier model got it right. Final weighted score: 93. Fourth place. Behind Kimi K3, GPT 5.6 Soul, and Claude Fable 5. Also, a tip if you want to try it: the personal coding plan link is basically unfindable on Alibaba Cloud. I had to ask their own AI agent to hand me the URL. That actually worked. CHAPTERS 0:00 Alibaba's huge claim about Qwen 3.8 Max 0:47 It's coding-plan only (and the signup link is hidden) 1:27 The nine tests on my benchmark, explained 2:58 Test 1: Halo Arena clone 4:30 Test 2: XCOM tactical grid 6:10 Test 3: 3D printable object (the hardest test) 7:24 Test 4: Sales page and animated SVGs 8:47 Coding, reasoning traps, and the skepticism audit 9:29 Grid Forge interactive spreadsheet 10:17 The final score and the leaderboard 12:01 Side by side vs Kimi K3 and Fable 5 15:34 Pricing, and should you actually use it? LINKS My site: https://mattjohnston.io I test every new model on this same benchmark the day it drops, so subscribe if you want the real numbers instead of the launch-day marketing.