OllamaLab
Flat isometric illustration of a dark chip package on a deep purple field, topped by a glowing blue crystal cube with an orange underglow and pink sparks
troubleshooting

Ollama Not Using GPU: How to Diagnose and Fix It

Ollama not using GPU: read the processor split, confirm the device is visible and supported, and fix the capacity or context setting that caused it.

By OllamaLab Editorial · ·Updated · 13 min read

If Ollama is not using your GPU, run ollama ps and read the PROCESSOR column. 100% CPU means the card was not detected: check the driver (NVIDIA 550 or newer), container GPU access, or an integrated GPU that Ollama skips by default. A CPU/GPU split means the model plus its context does not fit in free VRAM.

Ollama does not refuse to run a model that is too large for your GPU. It loads as many layers as fit and quietly runs the rest on the CPU. Nothing errors, nothing warns in the chat interface, and the only symptom is that generation is five to twenty times slower than expected. This is by far the most common cause of “my local model is slow”, and it is diagnosable in about thirty seconds.

How to tell if Ollama is using the GPU

Load a model, then run ollama ps. The PROCESSOR column is the definitive answer: 100% GPU means the model runs entirely on the GPU, a split such as 47%/53% CPU/GPU means only part of it does, and 100% CPU means the GPU is not being used at all. Reading that one line takes a few seconds and tells you both whether you have a problem and which kind. On Linux you can corroborate it in the same moment with nvidia-smi, which lists the ollama process holding VRAM and shows power draw and utilisation rising while a reply generates. If ollama ps says 100% GPU but nvidia-smi shows no Ollama process, you are looking at two different machines or two different users’ daemons.

Does Ollama use the GPU?

Yes, automatically, whenever it detects a supported card with enough free memory. Ollama offloads as many model layers to the GPU as fit and runs any remainder on the CPU, so in practice this is rarely a yes-or-no question: it uses as much of the GPU as the model and your free VRAM allow. The failure people actually hit is not Ollama refusing the GPU outright but Ollama using only part of it, which shows up as slow generation rather than an error message. Everything below is the path from that symptom to its cause, in the order that resolves it fastest.

Which GPUs does Ollama support?

Ollama supports NVIDIA cards with compute capability 5.0 or higher through CUDA, AMD Radeon and Instinct cards through ROCm v7, Apple Silicon through Metal, and a wider set of Windows and Linux GPUs through Vulkan, per the hardware support page. Integrated graphics are the exception: on Windows and Linux most are skipped unless you opt in.

HardwareBackendRequirement (Ollama docs, September 2026)
NVIDIA, compute 5.0+CUDADriver 550+ (570+ for compute 5.0 to 6.2); 551.61+ on Windows
AMD Radeon, Radeon PRO, InstinctROCmROCm v7 on Linux, HIP7-capable driver on Windows
Other Windows and Linux GPUsVulkanOn by default; OLLAMA_VULKAN=0 disables it
Most integrated graphicsVulkan or ROCmSkipped unless OLLAMA_IGPU_ENABLE=1 is set
Apple SiliconMetalBuilt in; see Apple Silicon memory sizing

The integrated-GPU row catches people out. Ollama’s GPU discovery code logs dropping integrated GPU; to enable, set OLLAMA_IGPU_ENABLE=1 and falls back to the CPU, and an Ollama update has done exactly that to a working AMD iGPU.

Step 1: read the split, before changing anything

With a model loaded, run:

ollama ps

The PROCESSOR column reports where the layers actually landed:

  • 100% GPU means full offload. Your GPU is being used, and slow generation here is a different problem, covered in Ollama tokens per second.
  • 47%/53% CPU/GPU or similar means partial offload. Part of every token is being computed at system-RAM speed.
  • 100% CPU means the GPU is not being used at all, which points at a detection or support problem rather than a capacity one. On a machine with no discrete GPU installed this is simply expected, and Ollama without a GPU covers what to expect from CPU-only inference instead.

The same output carries a CONTEXT column reporting the window actually allocated for the loaded model. That figure is worth reading alongside the split, because a larger context than you intended is a common reason a model that should fit does not.

Those three outcomes have different causes and different fixes, so establish which one you have before touching a setting. Skipping this step is how people spend an afternoon tuning sampling parameters on a model that was never on the GPU.

The server log carries the reasoning behind the decision. On Linux with a systemd service:

journalctl -u ollama --no-pager | tail -n 60

Look for the lines reporting detected devices, available memory, and the layer count offloaded. That log states what Ollama found and what it decided, which is more reliable than inference from symptoms.

Step 2: is the GPU visible at all?

If the split is 100% CPU, check whether the device was detected: the vendor tool has to list the card before Ollama can use it.

nvidia-smi          # NVIDIA
rocm-smi            # AMD with ROCm

If those commands fail or report nothing, the problem sits below Ollama and no Ollama setting will fix it. Driver installation, a reboot after a driver update, or a container missing GPU passthrough are the usual causes.

Common causes of an invisible GPU:

  • Driver not loaded after an update. A kernel or driver upgrade without a reboot leaves the userspace library and kernel module mismatched. nvidia-smi reports the mismatch directly.
  • Running in Docker without GPU access. The container needs the NVIDIA container toolkit and --gpus=all, or the equivalent device flags for AMD; how to run Ollama in Docker has both. Without them the container simply has no GPU. A container that cannot reach Ollama at all, rather than reaching it without a GPU, is a networking problem instead of an offload one: Open WebUI not connecting to Ollama covers that side, since a slow reply and a failed connection have entirely different fixes.
  • Unsupported compute capability, or a driver below the floor. The hardware support documentation gives two requirements, and the second is the one people miss: compute capability 5.0 or higher, and driver version 550 or newer, rising to 570 or newer for the oldest supported cards in the 5.0 to 6.2 range. A card that qualifies on paper will still be ignored on a distribution shipping an older driver branch. Check the version in the top-right of nvidia-smi output before concluding the card is unsupported.
  • CUDA_VISIBLE_DEVICES set to an empty or wrong value. A leftover export in a shell profile or systemd unit hides every device. Check the service environment as well as your shell.
  • The service runs as a different user. A systemd unit running as a user without access to the render or video group may not open the device even though your interactive shell can.

Step 3: it is detected, but the model does not fit

If nvidia-smi works and the split still shows CPU involvement, this is a capacity problem. The model plus its KV cache plus the compute buffers exceed free VRAM.

Check what is free, not what is installed:

nvidia-smi --query-gpu=memory.total,memory.used,memory.free --format=csv

A desktop session, a browser with hardware acceleration, or another loaded model can hold gigabytes before Ollama starts. On a machine with a display attached, that overhead is real and it is the difference between fitting and not fitting more often than people expect.

Then check whether the model should fit. At the Q4_K_M default, weights run about 0.6 GB per billion parameters, plus 1 to 3 GB for cache and buffers. The VRAM calculator will do that arithmetic for a specific model, quantisation, and context length, and how much VRAM you need to run local LLMs with Ollama explains where each term comes from.

Fixes, in the order that costs you the least quality:

  1. Close whatever else holds VRAM. Another resident model is the most common culprit. ollama ps lists them; ollama stop <model> unloads one.
  2. Reduce the context window. The KV cache grows linearly with context length, so this is often the cheapest large saving. Set num_ctx per request, or OLLAMA_CONTEXT_LENGTH server-wide.
  3. Quantise the KV cache. OLLAMA_KV_CACHE_TYPE=q8_0 roughly halves cache memory at negligible quality cost; q4_0 saves more and is likelier to show on long contexts. This requires flash attention to be active.
  4. Step down a model size. An 8B model fully on the GPU beats a 14B model half on the CPU, comfortably. On an 8 GB card in particular, the 14B class does not fit at Q4 by any arrangement of settings; best Ollama models for 8GB VRAM sets out what that tier does hold.
  5. Only then step down a quantisation level. Q4_K_M to Q3 is where quality loss becomes noticeable, particularly for structured output and tool calls.

Step 4: context length as a hidden cause

A model that loads cleanly can fall out of full GPU offload part-way through a long conversation, because the KV cache grows as the context fills. If ollama ps reported 100% GPU at load and generation slowed dramatically after a long document was pasted in, this is the mechanism.

Ollama’s default context window scales with available VRAM rather than being fixed, so the same model behaves differently on different machines. The context length documentation gives the tiers: 4k tokens below 24 GiB of VRAM, 32k from 24 to 48 GiB, and 256k at 48 GiB and above. Raising num_ctx to a model’s advertised maximum on a card that cannot hold the resulting cache is a reliable way to force partial offload.

Step 5: AMD-specific checks

AMD support runs through ROCm, with wider device coverage through Vulkan. The officially supported device list in the ROCm system requirements is narrower than AMD’s shipping product range, and unsupported-but-similar cards frequently need HSA_OVERRIDE_GFX_VERSION set to a supported architecture string, an x.y.z value such as 10.3.0 per the hardware support page, before ROCm will use them.

If ROCm cannot be made to work on a given card, the Vulkan backend is the fallback and covers a wider range of devices, generally at some performance cost. Confirm which backend is active from the server log rather than assuming.

Step 6: force Ollama to use the GPU with num_gpu

Ollama has no GPU-only switch; it already uses any supported GPU it can see. Forcing the GPU means making the device visible (Steps 2 and 5), then pinning the layer count with num_gpu.

Ollama estimates how many layers fit. The estimate is usually good, and occasionally conservative. To pin it:

curl http://localhost:11434/api/chat -d '{
  "model": "llama3.1:8b",
  "messages": [{"role": "user", "content": "hello"}],
  "options": { "num_gpu": 33, "num_ctx": 8192 }
}'

Inside ollama run, /set parameter num_gpu 33 does the same for the session, per the CLI help.

num_gpu sets the number of layers placed on the GPU. Setting it above what fits will cause an out-of-memory failure rather than a graceful fallback, which is arguably the more useful behaviour: an explicit error beats silent slowness.

Two environment variables are worth knowing:

  • OLLAMA_FLASH_ATTENTION=1 forces flash attention on, reducing attention memory as context grows. Ollama enables it automatically where supported, so this matters mainly when confirming it is active.
  • OLLAMA_SCHED_SPREAD=1 spreads a model across all available GPUs rather than fitting it onto the fewest. On a multi-GPU machine this trades some efficiency for capacity.

For multi-GPU systems where one card should be reserved, CUDA_VISIBLE_DEVICES restricts which devices Ollama sees. Set it in the service environment, not just your shell, or the daemon will not inherit it.

Why Ollama is not using the full GPU

A split such as 30%/70% CPU/GPU means the model plus its cache did not fit in free VRAM; Steps 3 and 4 fix that. Two other cases get described the same way:

  • Parallel requests. Per the FAQ, a 2K context serving 4 parallel requests allocates an 8K context, so raising OLLAMA_NUM_PARALLEL can tip a model into a split.
  • An idle second GPU. That is by design: a model that fits on one card stays on it. OLLAMA_SCHED_SPREAD=1 overrides that.

Utilisation below 100% while ollama ps reads 100% GPU is not one of them: every layer is on the card.

How to disable the GPU in Ollama

To disable the GPU for one request or session, set num_gpu to 0, which sends no layers to the card. To disable it for the whole server, start Ollama with CUDA_VISIBLE_DEVICES=-1 on NVIDIA or ROCR_VISIBLE_DEVICES=-1 on AMD; the hardware support page documents an invalid device ID as the way to force CPU.

In ollama run that is /set parameter num_gpu 0; in the API, "options": { "num_gpu": 0 }. Server-wide, add OLLAMA_VULKAN=0 so Vulkan does not step in. Ollama without a GPU covers that setup, and ollama ps should then read 100% CPU.

Ollama not using GPU on Windows

The Windows documentation requires Windows 10 22H2 or newer and NVIDIA driver 551.61 or newer. Past that:

  • Read server.log in %LOCALAPPDATA%\Ollama. Its discovery lines name the GPU and library chosen, and OLLAMA_DEBUG=1 adds detail, per the troubleshooting guide.
  • AMD through ROCm on Windows covers only the RX 7600 to 7900 XTX and Radeon PRO W7500 to W7900; RX 6000 cards may need Vulkan.
  • Hybrid graphics laptops. Issue #16667 had Ollama 0.30.7 drop an RTX 4080 Laptop GPU as “integrated” in favour of an Intel iGPU on Vulkan. It is closed with a linked pull request; update, or set OLLAMA_VULKAN=0.
  • Set variables under “Edit environment variables for your account” with Ollama quit from the taskbar, then restart it from the Start menu.

Ollama not using GPU on Ubuntu and other Linux

The official install script installs NVIDIA drivers through apt-get on Ubuntu and Debian, fetches the ROCm build for AMD cards, and adds the ollama user to the render and video groups. If it printed WARNING: No NVIDIA/AMD GPU detected. Ollama will run in CPU-only mode., get nvidia-smi or rocm-smi working and rerun it. A manual install needs the separate ollama-linux-amd64-rocm.tar.zst for AMD, and those groups added by hand. After a suspend cycle, reload NVIDIA’s UVM driver with sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm.

To force Ollama to use the GPU on Linux, variables must reach the systemd service:

  1. Run sudo systemctl edit ollama.service.
  2. Under [Service], add a line such as Environment="OLLAMA_IGPU_ENABLE=1".
  3. Run sudo systemctl daemon-reload, then sudo systemctl restart ollama.
  4. Reload the model and check ollama ps.

The diagnostic order that works

SymptomFirst checkLikely cause
100% CPU in ollama psnvidia-smi / rocm-smiDriver, container, or device support
100% CPU on a laptop or mini PCdropping integrated GPU in the logIntegrated GPU skipped by default
Partial split at loadFree VRAM vs model sizeModel plus cache exceeds capacity
Full GPU, then slows mid-chatnum_ctx and cache growthKV cache spilling as context fills
Worked yesterday, not todayDriver version, resident modelsUpdate without reboot, or memory held
100% GPU and still slowMemory bandwidth of the cardNot an offload problem at all

That last row matters. Once the split reads 100% GPU, offloading is no longer the issue and the speed you are getting is the speed the hardware provides. What sets that number is covered in Ollama tokens per second, and which cards deliver what is compared in best GPU for Ollama.

The general principle: verify the split first, verify capacity second, and change settings third. Reversing that order turns a thirty-second diagnosis into an afternoon.

FAQ

what does num_gpu do in ollama

num_gpu sets how many model layers Ollama sends to the GPU. The default, -1, lets Ollama decide dynamically when the model loads. Set 0 to keep a request on the CPU, or a fixed number to pin the split, through API options or /set parameter num_gpu in ollama run.

what does ollama need to use cuda

You need a working NVIDIA driver and a card with compute capability 5.0 or higher. Ollama’s hardware page sets the driver floor at 550, or 570 for compute 5.0 to 6.2 cards, and the Windows page asks for 551.61 or newer. On Ubuntu and Debian the official install script installs NVIDIA drivers itself; confirm with nvidia-smi.

can ollama use an integrated gpu

Yes, but on Windows and Linux Ollama skips most integrated GPUs by default and logs dropping integrated GPU; to enable, set OLLAMA_IGPU_ENABLE=1. Set that variable in the service environment and restart. Apple Silicon Macs use their GPU through Metal without it. Confirm with ollama ps.

why did ollama stop using my gpu after an update

Start with the driver. A kernel or NVIDIA driver update without a reboot leaves the module and library mismatched, and on Linux a suspend and resume can break the nvidia_uvm module; reload it with sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm. An Ollama update can also change which integrated GPUs it uses by default.

Sources

  1. Ollama Troubleshooting Guide
  2. Ollama Hardware Support
  3. Ollama FAQ
  4. AMD ROCm System Requirements
  5. Ollama Context Length
  6. Ollama on Windows
  7. Ollama on Linux
  8. Ollama source: discover/runner.go (integrated GPU filtering)
  9. Ollama source: api/types.go (num_gpu default)
  10. Ollama source: cmd/interactive.go (/set parameter help)
  11. Ollama Linux install script
  12. ollama/ollama issue #16667: Vulkan selects the Intel iGPU and drops the NVIDIA dGPU
  13. Ollama quietly stopped using my AMD iGPU

Related