Running Local LLMs in VS Code | Merge Conflict ep. 514

In this episode James and Frank dive into running AI coding models locally versus in the cloud—BYOK/Open Router, VS Code’s chat/agent harness, model runners (Olama, vLLM), and the practicality of 27B models on a 3090 using 4‑bit quantization. They share hands-on takeaways—how recent engineering (MT/MTPLX) boosts inference to usable token rates, when auto model selection makes sense, cost and hardware trade‑offs, and why local models can liberate your workflow while still needing smarter, unified tooling. Hosts: ‪@praeclarum‬ and ‪@JamesMontemagno‬ 00:00:00 - Intro & Guatemala Banter 00:04:00 - AI News & BYOK Intro 00:09:00 - Open Router & VS Code Models 00:16:00 - Running Local Models Setup 00:24:00 - Model Quality & Quantization 00:31:00 - Cloud vs Local Workflows 00:37:00 - Performance & Inference Speedups 00:44:00 - Hardware Limits, Future, Closing Merge Conflict: https://www.mergeconflict.fm Subscribe: https://www.mergeconflict.fm/subscribe Patreon:   / mergeconflictfm   #dotnet #podcast #mergeconflict