Running a 35B AI Model on 6GB VRAM, FAST (llama.cpp Guide)
Run a 35B parameter AI model on just 6GB VRAM using llama.cpp and Qwen 3.6. This setup shouldn’t work—but with the right optimizations, it reaches good enough tps on a GTX 1060. In this video, I break down how to run large language models locally on low VRAM GPUs using MoE offloading, memory tuning, and a few critical flags that dramatically improve performance. What you’ll learn: • How to run 35B LLMs on 6GB VRAM • llama.cpp optimization techniques • MoE (Mixture of Experts) offloading explained • Fixing slow token generation (3 tok/s → 17 tok/s) • Using --no-mmap and --mlock for performance and stability • TurboQuant for increasing context length • What doesn’t work (and why) Hardware used: • NVIDIA GTX 1060 (6GB VRAM) • Intel i3-8100 • 24GB RAM Tech stack: Proxmox → LXC → Docker → llama.cpp (adapt based on your setup) Useful resources: • Qwen 3.6 35B-A3B model: https://huggingface.co/Qwen/Qwen3.6-3... • TurboQuant paper: https://arxiv.org/abs/... • llama.cpp TurboQuant fork: https://github.com/TheTom/llama-cpp-t... If you're interested in running AI locally, optimizing LLM performance, or pushing old hardware to its limits, subscribe for more experiments. Chapters: 00:00 This shouldn’t work 00:27 Setup 01:46 Why it’s slow by default 02:52 MoE breakthrough 04:33 Fixing memory bottlenecks 05:32 Hitting 17 tok/s 06:40 4× context trick 09:23 Stability fix 11:04 What failed 13:32 The 5 flags #LocalAI #LLM #llamacpp #Qwen #AIonGPU #LowVRAM

192GB of VRAM in One PC… The Cheap Way

How do Graphics Cards Work? Exploring GPU Architecture

Can a 3.5GB model replace my 35B daily driver? (Bonsai 27B)

I Asked Claude Fable 5 to Improve llama.cpp.. and It Did

I tried Super Useful Future Tech! (That you can use Today)

Everything That Actually Matters for Local AI

Mojo + Vulkan is INSANE: Run Local AI on ANY GPU (Goodbye CUDA)

Bonsai 27B Runs Qwen 3.6 27B at 10x less memory.

Build Powerful Local Coding Agent on Budget GPU with Llama.cpp and Pi

The Non-NVIDIA AI Card Everyone’s Ignoring

I Tested the Cheapest Path to 96GB of VRAM

DFlash on GTX 1060: Can Dense AI Models Cheat VRAM Like MoE?

Why Google Just Gave Away Gemma 4 for Free

This Ridiculous $200 AI GPU Shouldn’t Be This Good

China quietly saved the world last month

Ultimate Guide Local AI Setup (Qwen3.6 + LlamaC++ + TurboQuant)

The Local AI Hardware Mistake Everyone Makes

The Best Local Agentic Coding Workflow (Complete Guide)

NVIDIA'S 748GB Ram Desktop Makes Local AI INSANELY Good

