If Ollama is not using your GPU, run ollama ps and read the PROCESSOR column. 100% CPU means the card was not detected: check the driver (NVIDIA 550 or newer), container GPU access, or an integrated GPU that Ollama skips by default. A CPU/GPU split means the model plus its context does not fit in free VRAM.
Ollama does not refuse to run a model that is too large for your GPU. It loads as many layers as fit and quietly runs the rest on the CPU. Nothing errors, nothing warns in the chat interface, and the only symptom is that generation is five to twenty times slower than expected. This is by far the most common cause of “my local model is slow”, and it is diagnosable in about thirty seconds.
How to tell if Ollama is using the GPU
Load a model, then run ollama ps. The PROCESSOR column is the definitive answer: 100% GPU means the model runs entirely on the GPU, a split such as 47%/53% CPU/GPU means only part of it does, and 100% CPU means the GPU is not being used at all. Reading that one line takes a few seconds and tells you both whether you have a problem and which kind. On Linux you can corroborate it in the same moment with nvidia-smi, which lists the ollama process holding VRAM and shows power draw and utilisation rising while a reply generates. If ollama ps says 100% GPU but nvidia-smi shows no Ollama process, you are looking at two different machines or two different users’ daemons.
Does Ollama use the GPU?
Yes, automatically, whenever it detects a supported card with enough free memory. Ollama offloads as many model layers to the GPU as fit and runs any remainder on the CPU, so in practice this is rarely a yes-or-no question: it uses as much of the GPU as the model and your free VRAM allow. The failure people actually hit is not Ollama refusing the GPU outright but Ollama using only part of it, which shows up as slow generation rather than an error message. Everything below is the path from that symptom to its cause, in the order that resolves it fastest.
Which GPUs does Ollama support?
Ollama supports NVIDIA cards with compute capability 5.0 or higher through CUDA, AMD Radeon and Instinct cards through ROCm v7, Apple Silicon through Metal, and a wider set of Windows and Linux GPUs through Vulkan, per the hardware support page. Integrated graphics are the exception: on Windows and Linux most are skipped unless you opt in.
| Hardware | Backend | Requirement (Ollama docs, September 2026) |
|---|---|---|
| NVIDIA, compute 5.0+ | CUDA | Driver 550+ (570+ for compute 5.0 to 6.2); 551.61+ on Windows |
| AMD Radeon, Radeon PRO, Instinct | ROCm | ROCm v7 on Linux, HIP7-capable driver on Windows |
| Other Windows and Linux GPUs | Vulkan | On by default; OLLAMA_VULKAN=0 disables it |
| Most integrated graphics | Vulkan or ROCm | Skipped unless OLLAMA_IGPU_ENABLE=1 is set |
| Apple Silicon | Metal | Built in; see Apple Silicon memory sizing |
The integrated-GPU row catches people out. Ollama’s GPU discovery code logs dropping integrated GPU; to enable, set OLLAMA_IGPU_ENABLE=1 and falls back to the CPU, and an Ollama update has done exactly that to a working AMD iGPU.
Step 1: read the split, before changing anything
With a model loaded, run:
ollama ps
The PROCESSOR column reports where the layers actually landed:
100% GPUmeans full offload. Your GPU is being used, and slow generation here is a different problem, covered in Ollama tokens per second.47%/53% CPU/GPUor similar means partial offload. Part of every token is being computed at system-RAM speed.100% CPUmeans the GPU is not being used at all, which points at a detection or support problem rather than a capacity one. On a machine with no discrete GPU installed this is simply expected, and Ollama without a GPU covers what to expect from CPU-only inference instead.
The same output carries a CONTEXT column reporting the window actually allocated for the loaded model. That figure is worth reading alongside the split, because a larger context than you intended is a common reason a model that should fit does not.
Those three outcomes have different causes and different fixes, so establish which one you have before touching a setting. Skipping this step is how people spend an afternoon tuning sampling parameters on a model that was never on the GPU.
The server log carries the reasoning behind the decision. On Linux with a systemd service:
journalctl -u ollama --no-pager | tail -n 60
Look for the lines reporting detected devices, available memory, and the layer count offloaded. That log states what Ollama found and what it decided, which is more reliable than inference from symptoms.
Step 2: is the GPU visible at all?
If the split is 100% CPU, check whether the device was detected: the vendor tool has to list the card before Ollama can use it.
nvidia-smi # NVIDIA
rocm-smi # AMD with ROCm
If those commands fail or report nothing, the problem sits below Ollama and no Ollama setting will fix it. Driver installation, a reboot after a driver update, or a container missing GPU passthrough are the usual causes.
Common causes of an invisible GPU:
- Driver not loaded after an update. A kernel or driver upgrade without a reboot leaves the userspace library and kernel module mismatched.
nvidia-smireports the mismatch directly. - Running in Docker without GPU access. The container needs the NVIDIA container toolkit and
--gpus=all, or the equivalent device flags for AMD; how to run Ollama in Docker has both. Without them the container simply has no GPU. A container that cannot reach Ollama at all, rather than reaching it without a GPU, is a networking problem instead of an offload one: Open WebUI not connecting to Ollama covers that side, since a slow reply and a failed connection have entirely different fixes. - Unsupported compute capability, or a driver below the floor. The hardware support documentation gives two requirements, and the second is the one people miss: compute capability 5.0 or higher, and driver version 550 or newer, rising to 570 or newer for the oldest supported cards in the 5.0 to 6.2 range. A card that qualifies on paper will still be ignored on a distribution shipping an older driver branch. Check the version in the top-right of
nvidia-smioutput before concluding the card is unsupported. CUDA_VISIBLE_DEVICESset to an empty or wrong value. A leftover export in a shell profile or systemd unit hides every device. Check the service environment as well as your shell.- The service runs as a different user. A systemd unit running as a user without access to the render or video group may not open the device even though your interactive shell can.
Step 3: it is detected, but the model does not fit
If nvidia-smi works and the split still shows CPU involvement, this is a capacity problem. The model plus its KV cache plus the compute buffers exceed free VRAM.
Check what is free, not what is installed:
nvidia-smi --query-gpu=memory.total,memory.used,memory.free --format=csv
A desktop session, a browser with hardware acceleration, or another loaded model can hold gigabytes before Ollama starts. On a machine with a display attached, that overhead is real and it is the difference between fitting and not fitting more often than people expect.
Then check whether the model should fit. At the Q4_K_M default, weights run about 0.6 GB per billion parameters, plus 1 to 3 GB for cache and buffers. The VRAM calculator will do that arithmetic for a specific model, quantisation, and context length, and how much VRAM you need to run local LLMs with Ollama explains where each term comes from.
Fixes, in the order that costs you the least quality:
- Close whatever else holds VRAM. Another resident model is the most common culprit.
ollama pslists them;ollama stop <model>unloads one. - Reduce the context window. The KV cache grows linearly with context length, so this is often the cheapest large saving. Set
num_ctxper request, orOLLAMA_CONTEXT_LENGTHserver-wide. - Quantise the KV cache.
OLLAMA_KV_CACHE_TYPE=q8_0roughly halves cache memory at negligible quality cost;q4_0saves more and is likelier to show on long contexts. This requires flash attention to be active. - Step down a model size. An 8B model fully on the GPU beats a 14B model half on the CPU, comfortably. On an 8 GB card in particular, the 14B class does not fit at Q4 by any arrangement of settings; best Ollama models for 8GB VRAM sets out what that tier does hold.
- Only then step down a quantisation level. Q4_K_M to Q3 is where quality loss becomes noticeable, particularly for structured output and tool calls.
Step 4: context length as a hidden cause
A model that loads cleanly can fall out of full GPU offload part-way through a long conversation, because the KV cache grows as the context fills. If ollama ps reported 100% GPU at load and generation slowed dramatically after a long document was pasted in, this is the mechanism.
Ollama’s default context window scales with available VRAM rather than being fixed, so the same model behaves differently on different machines. The context length documentation gives the tiers: 4k tokens below 24 GiB of VRAM, 32k from 24 to 48 GiB, and 256k at 48 GiB and above. Raising num_ctx to a model’s advertised maximum on a card that cannot hold the resulting cache is a reliable way to force partial offload.
Step 5: AMD-specific checks
AMD support runs through ROCm, with wider device coverage through Vulkan. The officially supported device list in the ROCm system requirements is narrower than AMD’s shipping product range, and unsupported-but-similar cards frequently need HSA_OVERRIDE_GFX_VERSION set to a supported architecture string, an x.y.z value such as 10.3.0 per the hardware support page, before ROCm will use them.
If ROCm cannot be made to work on a given card, the Vulkan backend is the fallback and covers a wider range of devices, generally at some performance cost. Confirm which backend is active from the server log rather than assuming.
Step 6: force Ollama to use the GPU with num_gpu
Ollama has no GPU-only switch; it already uses any supported GPU it can see. Forcing the GPU means making the device visible (Steps 2 and 5), then pinning the layer count with num_gpu.
Ollama estimates how many layers fit. The estimate is usually good, and occasionally conservative. To pin it:
curl http://localhost:11434/api/chat -d '{
"model": "llama3.1:8b",
"messages": [{"role": "user", "content": "hello"}],
"options": { "num_gpu": 33, "num_ctx": 8192 }
}'
Inside ollama run, /set parameter num_gpu 33 does the same for the session, per the CLI help.
num_gpu sets the number of layers placed on the GPU. Setting it above what fits will cause an out-of-memory failure rather than a graceful fallback, which is arguably the more useful behaviour: an explicit error beats silent slowness.
Two environment variables are worth knowing:
OLLAMA_FLASH_ATTENTION=1forces flash attention on, reducing attention memory as context grows. Ollama enables it automatically where supported, so this matters mainly when confirming it is active.OLLAMA_SCHED_SPREAD=1spreads a model across all available GPUs rather than fitting it onto the fewest. On a multi-GPU machine this trades some efficiency for capacity.
For multi-GPU systems where one card should be reserved, CUDA_VISIBLE_DEVICES restricts which devices Ollama sees. Set it in the service environment, not just your shell, or the daemon will not inherit it.
Why Ollama is not using the full GPU
A split such as 30%/70% CPU/GPU means the model plus its cache did not fit in free VRAM; Steps 3 and 4 fix that. Two other cases get described the same way:
- Parallel requests. Per the FAQ, a 2K context serving 4 parallel requests allocates an 8K context, so raising
OLLAMA_NUM_PARALLELcan tip a model into a split. - An idle second GPU. That is by design: a model that fits on one card stays on it.
OLLAMA_SCHED_SPREAD=1overrides that.
Utilisation below 100% while ollama ps reads 100% GPU is not one of them: every layer is on the card.
How to disable the GPU in Ollama
To disable the GPU for one request or session, set num_gpu to 0, which sends no layers to the card. To disable it for the whole server, start Ollama with CUDA_VISIBLE_DEVICES=-1 on NVIDIA or ROCR_VISIBLE_DEVICES=-1 on AMD; the hardware support page documents an invalid device ID as the way to force CPU.
In ollama run that is /set parameter num_gpu 0; in the API, "options": { "num_gpu": 0 }. Server-wide, add OLLAMA_VULKAN=0 so Vulkan does not step in. Ollama without a GPU covers that setup, and ollama ps should then read 100% CPU.
Ollama not using GPU on Windows
The Windows documentation requires Windows 10 22H2 or newer and NVIDIA driver 551.61 or newer. Past that:
- Read
server.login%LOCALAPPDATA%\Ollama. Its discovery lines name the GPU and library chosen, andOLLAMA_DEBUG=1adds detail, per the troubleshooting guide. - AMD through ROCm on Windows covers only the RX 7600 to 7900 XTX and Radeon PRO W7500 to W7900; RX 6000 cards may need Vulkan.
- Hybrid graphics laptops. Issue #16667 had Ollama 0.30.7 drop an RTX 4080 Laptop GPU as “integrated” in favour of an Intel iGPU on Vulkan. It is closed with a linked pull request; update, or set
OLLAMA_VULKAN=0. - Set variables under “Edit environment variables for your account” with Ollama quit from the taskbar, then restart it from the Start menu.
Ollama not using GPU on Ubuntu and other Linux
The official install script installs NVIDIA drivers through apt-get on Ubuntu and Debian, fetches the ROCm build for AMD cards, and adds the ollama user to the render and video groups. If it printed WARNING: No NVIDIA/AMD GPU detected. Ollama will run in CPU-only mode., get nvidia-smi or rocm-smi working and rerun it. A manual install needs the separate ollama-linux-amd64-rocm.tar.zst for AMD, and those groups added by hand. After a suspend cycle, reload NVIDIA’s UVM driver with sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm.
To force Ollama to use the GPU on Linux, variables must reach the systemd service:
- Run
sudo systemctl edit ollama.service. - Under
[Service], add a line such asEnvironment="OLLAMA_IGPU_ENABLE=1". - Run
sudo systemctl daemon-reload, thensudo systemctl restart ollama. - Reload the model and check
ollama ps.
The diagnostic order that works
| Symptom | First check | Likely cause |
|---|---|---|
100% CPU in ollama ps | nvidia-smi / rocm-smi | Driver, container, or device support |
100% CPU on a laptop or mini PC | dropping integrated GPU in the log | Integrated GPU skipped by default |
| Partial split at load | Free VRAM vs model size | Model plus cache exceeds capacity |
| Full GPU, then slows mid-chat | num_ctx and cache growth | KV cache spilling as context fills |
| Worked yesterday, not today | Driver version, resident models | Update without reboot, or memory held |
100% GPU and still slow | Memory bandwidth of the card | Not an offload problem at all |
That last row matters. Once the split reads 100% GPU, offloading is no longer the issue and the speed you are getting is the speed the hardware provides. What sets that number is covered in Ollama tokens per second, and which cards deliver what is compared in best GPU for Ollama.
The general principle: verify the split first, verify capacity second, and change settings third. Reversing that order turns a thirty-second diagnosis into an afternoon.
FAQ
what does num_gpu do in ollama
num_gpu sets how many model layers Ollama sends to the GPU. The default, -1, lets Ollama decide dynamically when the model loads. Set 0 to keep a request on the CPU, or a fixed number to pin the split, through API options or /set parameter num_gpu in ollama run.
what does ollama need to use cuda
You need a working NVIDIA driver and a card with compute capability 5.0 or higher. Ollama’s hardware page sets the driver floor at 550, or 570 for compute 5.0 to 6.2 cards, and the Windows page asks for 551.61 or newer. On Ubuntu and Debian the official install script installs NVIDIA drivers itself; confirm with nvidia-smi.
can ollama use an integrated gpu
Yes, but on Windows and Linux Ollama skips most integrated GPUs by default and logs dropping integrated GPU; to enable, set OLLAMA_IGPU_ENABLE=1. Set that variable in the service environment and restart. Apple Silicon Macs use their GPU through Metal without it. Confirm with ollama ps.
why did ollama stop using my gpu after an update
Start with the driver. A kernel or NVIDIA driver update without a reboot leaves the module and library mismatched, and on Linux a suspend and resume can break the nvidia_uvm module; reload it with sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm. An Ollama update can also change which integrated GPUs it uses by default.