QWEN3.6-35B-A3B-MTP + VISION: DOES --MMPROJ KILL THE MTP DRAFT?
I've seen maybe some problems in the past with this model, so maybe an update it is welcome to all thos ewho use it. The concern: https://www.reddit.com/r/llamacpp/comments... The answer: No — --mmproj does not kill the MTP draft. Not on the vision turn, and not for the rest of the session. llama.cpp b9620 · Qwen3.6-35B-A3B-MTP · M5 Max · unedited. Everything is on screen: the server log on the left, the model answering on the right. Watch 2:14 — the image description lands while the same log prints "draft acceptance = 0.61950" at 104.68 t/s. One frame, both facts. THE LAUNCH llama-server -hf unsloth/Qwen3.6-35B-A3B-MTP-GGUF:UD-Q4_K_XL \ --mmproj .../mmproj-BF16.gguf \ -ngl 99 -c 262144 -fa on -np 1 \ --spec-type draft-mtp --spec-draft-n-max 2 --port 8081 Both budgets appear side by side at 0:38 — the projector does not evict the draft context: [mtmd] estimated worst-case memory usage of mmproj is 1134.00 MiB [spec] estimated memory usage of MTP context is 826.70 MiB common_speculative_impl_draft_mtp: adding speculative implementation 'draft-mtp' CHAPTERS 0:00 CO_DE, opening the hajimi project 0:20 New desk → new terminal, side by side with the chat 0:30 Launching the MTP + mmproj server 0:38 Boot log: mmproj AND MTP draft both initialize 0:45 TEXT test — "quick audit of this repo's entry points" 1:14 Drafting on text: 0.862 → 0.955 acceptance 1:50 The audit result (real tables, local model) 2:00 VISION test — "describe this image" 2:02 process_mtmd: encoding mtmd batch 2:14 The answer + draft acceptance 0.6195 @ 104.68 t/s 2:20 Conclusion THE NUMBERS Text turns 0.862 / 0.955 / 0.828 / 0.739 acceptance · 79–100 t/s Vision turn 0.6195 (394 accepted / 636 generated) · 104.68 t/s — fastest eval of the whole session After vision 0.680 — a text turn AFTER the image still drafts Session total 1474 / 2020 draft tokens accepted (~73%) across 7 turns Acceptance does drop on the vision turn. That is the draft head being less sure about tokens that follow an image — lower acceptance, not a disabled draft. Throughput did not suffer. If you do OCR/grounding, the loader warns: "Qwen-VL models require at minimum 1024 image tokens … try --image-min-tokens 1024" THE WORKSPACE The app is CO_DE — a native macOS workspace where CLI coding agents (Claude Code, Codex, Cursor, Gemini) sit next to your terminals, with a local models in CLIs like PI, Opencode, or in native chat the orchestrator seat. That's why the server log and the model answering it are on one screen: same machine, same session, nothing spliced. CO_DE: https://github.com/gelubodrug/co_de Reddit thread: https://www.reddit.com/r/llamacpp/comments...

The PROBLEM with Capitalism - Smarter Every Day 316

Jon Stewart on Trump's "Meritocracy" & Desi Lydic on MAGA Moving Iran Goalposts | The Daily Show

VRAM vs. RAM vs. SSD | Where Do Local AI Models Really Run?

Fall asleep while I build a town into a city (from nothing)

I Tried Coding my own Graphics Library

Tested ik_llama.cpp vs llama.cpp: Qwen3.6-35B-A3B-MTP Text + Vision

Chichvarkin Saves Putin | Vitaly Portnikov

Комиссаренко, Белый, Детков «Давайте вместе сделаем шоу #8»

AMD Says 2 Ryzen AI Halos Can Run a 400B Model... I Tested It

This Bassett Hound WAS Trying to Tell me Something…He Was Right | PUPDATE

China quietly saved the world last month

Casey Muratori – The Big OOPs: Anatomy of a Thirty-five-year Mistake – BSC 2025

🐧 Linux Web Server - Installing and configuring FTP with VSFTPD on Linux CentOS - Lesson 6 Part 1/2

No Boss, No Money: The Raw Reality of China’s Gen-Z Freelancers

System Design Course – APIs, Databases, Caching, CDNs, Load Balancing & Production Infra

777 Portal ✨ Manifest Miracles, Abundance & Divine Alignment - Meditation Music

System Design Explained: APIs, Databases, Caching, CDNs, Load Balancing & Production Infra

Is This Wish Meant to Be Fulfilled? 🧚🤲 Detailed Pick a Card Tarot Reading ✫・

Ich teste AMAZON BESTSELLER 💸 (worth the hype oder nicht?)

