OllamaLab
Industrial AI Hardware Spec Tool

Ollama Local LLM VRAM & Speed Calculator

Estimate exact GPU memory requirements, KV cache allocation, quantization overhead, and expected tokens/second throughput across NVIDIA, AMD, and Apple Silicon hardware.

Model & Quantization Spec

16k

GPU & Memory Subsystem

Total VRAM Required
41.8 GB
Weights + KV Cache
Expected Inference Speed
24 t/s
Single Batch Generation
GPU Fit Assessment
DUAL GPU
Requires 2x 24GB VRAM

VRAM Allocation Breakdown Calculated Specs

Memory Component VRAM Overhead
Model Weights (Quantized) 39.4 GB
KV Cache (Context Storage) 1.4 GB
CUDA Context & Buffer Overhead 1.0 GB
Total Target VRAM Needed 41.8 GB
Ollama Run Command
ollama run llama3:70b-instruct-q4_K_M

Where these numbers come from

The memory total is weights plus KV cache plus buffers. Weights are parameter count multiplied by bits per weight, which is what the quantisation setting changes. The speed estimate divides a card's memory bandwidth by the bytes read per generated token, because single-stream generation is bound by memory bandwidth rather than by compute. Both figures are estimates from published specifications, so treat them as a sizing guide rather than a measurement of your machine.

Each guide below works through one part of that arithmetic: