Edge ai

Blog posts tagged “Edge AI”


Small LLMs That Fit in 8GB: The Best Models to Self-Host in 2026

Local LLM Self-Hosted AI Ollama

Which open-weight LLMs actually fit in 8GB of VRAM or RAM in 2026, with measured file sizes, KV cache math from published configs, and Ollama commands for Qwen3.5, Gemma 4, Ministral 3, Granite 4.1, Nemotron 3 Nano, and Phi-4-mini.

TurboQuant for Efficient LLMs and How Gemma 4 Utilizes It

TurboQuant Gemma 4 On-Device AI

Learn what TurboQuant is, the math behind Google's new compression method, and how Gemma 4 combines efficient architectures and edge runtimes to run on phones and other edge devices.