LLM throughput benchmark on 13 GPUs and 10 models
LLM throughput benchmark on 13 GPUs and 10 models. Article with data and commands: https://pavlokhmel.com/llm_throughput... Q4_K_M LLM models unsloth/LFM2.5-1.2B-Instruct-GGUF unsloth/Ministral-3-3B-Instruct-2512-GGUF unsloth/Qwen3.5-4B-GGUF unsloth/gemma-4-E4B-it-GGUF unsloth/Qwen3.5-9B-GGUF unsloth/gpt-oss-20b-GGUF unsloth/Qwen3.6-27B-GGUF unsloth/gemma-4-26B-A4B-it-GGUF unsloth/gemma-4-31B-it-GGUF unsloth/Qwen3.6-35B-A3B-GGUF GPUs Macmini9,1 - M1 - 8 cores (4 Performance and 4 Efficiency) - 16GB RAM MacBook Pro - Mac17,2 - 10 cores (4 Super and 6 Efficiency) - 32GB RAM NVIDIA Tesla P100 PCIe 16 GB NVIDIA GeForce GTX 1080 8GB NVIDIA Tesla V100 PCIe 32 GB NVIDIA A100 PCIe 40 GB NVIDIA A100 SXM 80 GB NVIDIA GeForce RTX 4080 16GB NVIDIA GeForce RTX 4090 24GB NVIDIA H200 NVL NVIDIA H100 SXM 80GB NVIDIA H200 SXM 141GB NVIDIA B300 SXM 288GB

This “GPU” Has Network Ports… WTF

192GB of VRAM in One PC… The Cheap Way

The Scariest Chart in Electrical Engineering

The Best Local Agentic Coding Workflow (Complete Guide)

Up to 8x Faster AI N-gram Explained, Deployed & Benchmarked on Qwen 3 6 27B Lamma cpp!

$1500 Local AI Server Build Tested with Hermes Agent Gemma 4 and Qwen 3.6

DO NOT BUY: LG’s Spyware TVs, Monitors, and Wiretapping Concerns

NVIDIA didn't want me to do this

The Invisible Phone: Ditch PSTN Forever – No IMEI, No IMSI, No KYC, Total Stealth Calling

Mojo + Vulkan is INSANE: Run Local AI on ANY GPU (Goodbye CUDA)

Opus 5 is INSANE...

this Raspberry Pi belongs on your wall

Up to 6x Faster AI? DFlash Explained, Deployed & Benchmarked on Qwen 3.6 27B. Lamma.cpp!

AI Isn't Replacing Your Job. That's the Problem.

No Boss, No Money: The Raw Reality of China’s Gen-Z Freelancers

AMD Just Killed AI Subscriptions Forever (Ryzen AI Halo)

Fall asleep while I build a town (from nothing) - "Town To City" ASMR for sleep

Little Coder: PI-based harness for small models

Italian Deep House 2026 | Mediterranean Vibes

