Topics
Browse OllamaLab by topic: VRAM sizing, GPU memory bandwidth, quantisation, CPU-only inference and Ollama troubleshooting, newest posts first.
Tags
- #local-llm 9
- #ollama 9
- #gpu 6
- #inference 5
- #hardware 4
- #memory-bandwidth 4
- #vram 4
- #performance 3
- #quantization 3
- #rocm 2
- #apple-silicon 1
- #comparison 1
- #cpu-inference 1
- #cuda 1
- #docker 1
- #llama-cpp 1
- #lm-studio 1
- #mlx 1
- #nvidia 1
- #troubleshooting 1
Categories
hardware 4 posts
- Best GPU for Ollama: 8GB to 48GB Cards ComparedBest GPU for Ollama by the two specs that matter: VRAM capacity decides what loads, memory bandwidth decides how fast it answers. Compared by tier.
- Ollama on Apple Silicon: Unified Memory SizingOllama on Apple Silicon: how much unified memory a model actually gets, what each M-series chip's bandwidth means for speed, and which Mac to size for.
- Best Ollama Models for 8GB VRAM: What FitsBest Ollama models for 8GB VRAM: which parameter counts and quantisations fit once the desktop takes its share, and what to run at each context length.
- How Much VRAM Do You Need to Run Local LLMs with Ollama?VRAM sizing for local LLMs: how quantisation, parameter count, and context length set your GPU memory bill, and which model fits 8, 12, 16, or 24 GB.
performance 2 posts
- Ollama Tokens per Second: What Sets Your SpeedOllama tokens per second explained: how memory bandwidth, quantisation, and context length set generation rate, plus how to measure your own numbers.
- Ollama Without a GPU: CPU-Only Speed and RAMOllama without a GPU: what CPU-only inference costs in tokens per second, how much system RAM each model class needs, and when it is still worth running.