Local LLM performance tuning
Let's take a look at our dual-GPU local LLM setup and see if we can tune the settings to squeeze a few more tokens per second out of it! We'll talk about the advantages and disadvantages of MTP (multi-token prediction), and demonstrate what changing a few startup options to llama-server can do! 00:00 Intro 00:49 Baseline reading 01:43 Layer split 03:28 Tensor split 04:16 Prompt prefill vs generation 05:02 MTP 06:46 Enabling MTP 08:01 What's the catch? 09:55 The tradeoff 10:34 Conclusion

▶︎
LLM Battle 2 - Qwen, Ornith, Qwythos, Gemma 4

▶︎
Qwen 3.6 14B A3B FableVibes benchmarked and tested vs Base Qwen 35B - 16GB Local LLM setup

▶︎
Building a completely local AI code assistant

▶︎
Day 9: 100 Days of Genai for DevOps: DevOps Interview Buddy using n8n, Ollama, and Gemma(Hindi)

▶︎
192GB of VRAM in One PC… The Cheap Way

▶︎
Building llama.cpp from source

▶︎
The RAM Price CRASH Will Be Apocalyptic!

▶︎
Kimi K3 explained in 13min..

▶︎
Running a local LLM with two GPUs

▶︎
The Non-NVIDIA AI Card Everyone’s Ignoring

▶︎
I Bought an NVIDIA DGX Spark... Here's What It Actually Does (90+ Tokens/sec)

▶︎
Building agents and skills with OpenCode

▶︎
I Didn’t Expect Local AI to Go This Far on a Laptop

▶︎
Train Your Own Tiny LLM on Your PC In Just A Few Hours

▶︎
I Spent 1.5 Billion AI Tokens on Absolute Slop

▶︎
Little Coder: Small Models Need Small Harnesses

▶︎
I Asked Claude Fable 5 to Improve llama.cpp.. and It Did

▶︎
Configuring OpenCode for a local LLM

▶︎
The Most Important Conversation in AI Right Now

▶︎
