About OllamaLab
OllamaLab covers the hardware side of running local models with Ollama. Two numbers decide almost everything: how much VRAM a given model and quantisation actually needs, and how much memory bandwidth the card has to read those weights back on every generated token. The rest of the coverage works outward from those two, through context-length cost, partial CPU offload, PCIe bandwidth, and what happens when a model spills into system RAM.
It is written for people sizing a GPU before they buy one, and for people working out why the card they already own is slower than they expected. Memory and throughput figures are taken from vendor specifications, model cards, Ollama and llama.cpp documentation, and published measurements, each linked at the foot of the article it appears in.
What is covered here
- Comparison
- Hardware
- Performance
- Setup
- Troubleshooting
9 articles are published so far. New articles are announced on the RSS feed; there is no fixed publishing schedule and this site does not promise one.
The calculator
The VRAM calculator is the one piece of this site that is not an article. Give it a model size, a quantisation level, a context window, and a card, and it returns the memory total the model will need and an estimated generation rate for that card's memory bandwidth. It applies the same arithmetic the articles explain, so the two are meant to be used together rather than separately.
How these articles are produced
Articles are researched from primary sources: vendor and project documentation, published standards and specifications, release notes, advisories, and measurements published by the people who took them. Drafts are produced with AI assistance and then edited against those same sources before anything is published. Where a figure comes from a datasheet or a third-party measurement, the article names the source and links to it so you can check the original rather than take this site's summary of it.
Everything here is published under the OllamaLab Editorial byline. That is an editorial desk, not a person, and no article on this site claims hands-on lab testing, benchmarking, or first-hand measurement. Nothing here should be read as a report of something this site physically tested.
Corrections
Getting it right matters more than getting it first. If something on this site is wrong, out of date, or missing the source it should cite, email hello@ollamalab.com with the page and the specific claim. Substantive corrections are made on the page itself rather than quietly dropped.
How this site is funded
This site currently runs no affiliate links, no sponsored content, no paid placement, and no display advertising. Nothing on it earns a commission. If that changes, this page and the disclosure page will say so before any such link appears.
The full position is on the disclosure page. Read it before acting on anything here that reads like a buying recommendation.
Related sites
OllamaLab is run alongside a small number of other single-topic sites:
- LLMOps Report - Operating LLMs in production — eval, observability, cost, latency.
- ProxmoxGuide - Virtualize with confidence.
Contact
Email: hello@ollamalab.com
Site: ollamalab.com
Privacy: privacy policy ·
terms of use