OllamaLab

OllamaLab — Local LLM hardware and inference speed.

Local LLM hardware and inference speed.
OllamaLab
Independent · 2026
Isometric dark navy scene of a compact PC with a glowing processor cabled through a cyan chip to a laptop, evoking local LLM runners on personal hardware Featured
comparison

Ollama vs LM Studio: Which Local LLM Runner to Pick

Ollama vs LM Studio compared on speed, API, headless serving, Apple Silicon MLX and licensing, with a clear pick for developers and desktop users.

Read the article →

Latest reviews

Start here: sizing and speed for local models

OllamaLab covers the hardware side of running local models with Ollama: how much GPU memory a model needs, how fast a card can read those weights back on every generated token, and what to change when a model runs slower than it should. Two numbers decide almost everything - the VRAM a model and its quantisation occupy, and the memory bandwidth that sets generation speed once it fits. Every guide below works outward from those two.

Work out what fits

Make it fast, and fix it when it is slow

Specific setups