Ollama Not Using GPU? Why It Runs on CPU and How to Fix It (NVIDIA, AMD, Intel)

Logeshwaran
—

Ollama uses your GPU automatically when it finds a supported one, so when it runs on the CPU instead, one of a few things is wrong: the model (or its context) does not fit in your graphics memory, the GPU driver is too old or missing, your card is not on Ollama's supported list, or an environment variable is hiding the GPU. Here is the surprise that sends most people down the wrong path: on Windows, Task Manager can show your GPU at 1% while Ollama is running on it, because the default graph shows 3D graphics work, not AI compute. The real answer takes ten seconds. Start a model, then run ollama ps in a terminal and read the PROCESSOR column: 100% GPU is good, 100% CPU means no GPU was used, and a split like 48%/52% CPU/GPU means the model did not fit. Below, we start with the basics, then fix every case: NVIDIA on Windows and Linux, AMD Radeon, Intel Arc, Docker, WSL2 and Mac.

Jake runs a small phone repair shop, and he wanted a local AI model to draft repair quotes and replies to customers, without sending anyone's name or phone number to a cloud service. He bought a used graphics card with 12 GB of memory for $180, installed Ollama, downloaded a 14-billion-parameter model a video recommended, and set the context to 32,000 tokens because the same video said bigger was better. Each reply took about 40 seconds. He opened Task Manager, saw the GPU at 1%, and decided Ollama was ignoring the card he had just paid for. He spent an evening reinstalling drivers and Ollama twice. Ethan looked at it the next morning, typed one command, and pointed at a single column. Ollama was using the GPU the whole time. It just could not fit everything on it.

⚡ Quick Answer

• Check first → run ollama ps while a model is loaded. Read the PROCESSOR column.

• Split CPU/GPU → the model or context is too big for your VRAM. Make it fit.

• 100% CPU with NVIDIA → driver 550 or newer. Windows or Linux.

• AMD or Intel GPU ignored → AMD or Intel, and update Ollama.

Docker or WSL2? Docker and WSL2 have their own fixes.

🧭 NEW HERE? READ THESE FIRST

New to running AI models locally? These five pages cover the basics this one builds on:

📌 Bookmark this; ollama ps first, the server log second, drivers third.

The basics: why a GPU makes Ollama fast, and what VRAM is

Before fixing anything, it helps to know what Ollama is actually doing with your hardware. Four ideas cover it.

  • A model is a big file of numbers. When Ollama "loads" a model, it copies that file into memory, then does an enormous amount of simple math on it for every word it writes.
  • A GPU is built for exactly that math. A graphics card has thousands of small cores that do the same calculation side by side, so it can produce text many times faster than a CPU.
  • VRAM is the GPU's own memory. The GPU can only work fast on what sits in its VRAM. Your normal system memory (RAM) is a separate pool that the CPU uses. A card with 8 GB of VRAM can hold about 8 GB of model, minus a little.
  • Layers can be split. A model is made of layers, stacked one after another. If the whole model does not fit in VRAM, Ollama puts as many layers as it can on the GPU and runs the rest on the CPU. That is called offloading, and it is why you sometimes see a split instead of a clean "GPU" or "CPU."

"So it's like a kitchen with a fast chef and a slow one," Jake said.

"Close," Ethan said. "The fast chef has a small counter. Whatever doesn't fit on his counter goes to the slow chef, and the meal comes out at the slow chef's speed. Your card's counter is 12 GB. You handed him about 14."

CUDA, ROCm, Vulkan and Metal: the four ways Ollama talks to a GPU

A GPU needs a "backend," a software bridge from the model to the card. Ollama picks one automatically:

BackendUsed forWhat it needs from you
CUDANVIDIA GeForce, RTX, Quadro and data center cardsAn up-to-date NVIDIA driver; Ollama brings the rest
ROCmAMD Radeon and Instinct cards on its supported listA supported card and a current AMD driver
VulkanIntel Arc and Iris Xe, and AMD cards outside the ROCm listA current graphics driver and a current Ollama
MetalApple Silicon Macs (M1 and later)Nothing; it is automatic

You never install CUDA or ROCm separately just for Ollama; it ships with the libraries it needs. What you do need is the right driver for your card, and a card Ollama recognizes. Most of the fixes below come down to one of those two.

Does Ollama automatically use the GPU?

Yes. There is no "use GPU" switch to turn on. When the Ollama server starts, it looks for supported GPUs, and when you run a model, it fills the GPU first and only falls back to the CPU for whatever does not fit, or for everything if it found no usable GPU. So the question is never "how do I turn the GPU on?" It is "why did Ollama decide it could not use it?" The next two sections show you where Ollama answers that question.

Does Ollama run on GPU or CPU? How to check in 30 seconds

  1. Start a model in one terminal, for example ollama run llama3.2, and ask it anything so it loads.
  2. Open a second terminal and run ollama ps.
  3. Read the PROCESSOR column for your model.
PROCESSOR saysWhat it meansGo to
100% GPUThe whole model is in VRAM. Ollama is using your GPU fully.Nothing to fix; if it still feels slow, see here
48%/52% CPU/GPU (any split)The GPU is working, but part of the model did not fit, and that part runs on the CPU.Make it fit
100% CPUOllama found no GPU it could use, or something told it not to.Read the log
Nothing listedNo model is loaded right now. Models unload after five idle minutes by default.Run a prompt, then check again

Ethan ran it on Jake's PC. The line read 31%/69% CPU/GPU. "There," he said. "Sixty-nine percent of the model is on your new card. Thirty-one percent is on the CPU, and that slice is setting the speed of every reply."

The Task Manager trap on Windows

This is where Jake went wrong, and he is far from alone. Open Task Manager, go to Performance and click your GPU. The big graphs show 3D, Copy, Video Encode and Video Decode by default. AI work does not run on the 3D engine, so those graphs can sit near zero while Ollama is busy on the card.

Two ways to see the truth in Task Manager:

  • Watch the memory, not the graphs. Dedicated GPU memory jumps by several gigabytes the moment a model loads onto the GPU. If it barely moves when you load a model, the model is not on the GPU.
  • Switch a graph to Cuda or Compute. Click the small name above one of the graphs (for example "Video Encode") and pick Cuda or Compute_0 from the list. That graph now shows AI work. On some PCs with hardware-accelerated GPU scheduling turned on, the Cuda option is not offered; use the memory reading or nvidia-smi instead.

Other ways to check: nvidia-smi, AMD and Intel tools, Mac

On any PC with an NVIDIA card, Windows or Linux, open a terminal and run nvidia-smi. While a model is loaded, the bottom table lists an Ollama process and how much GPU memory it uses, and the top shows the card's total memory. Run nvidia-smi -l 1 to refresh every second while the model answers; GPU utilization should jump with every reply. On Linux with AMD cards, amd-smi or rocm-smi shows the same thing, and for Intel graphics on Linux, intel_gpu_top (from the intel-gpu-tools package) does. On a Mac, open Activity Monitor and choose Window → GPU History.

If you are not sure which graphics card your PC even has, the dxdiag tool shows the exact model and its memory under the Display tab.

Read the server log: the line that tells you why

ollama ps tells you what happened. The server log tells you why. Every time the Ollama server starts, it writes down which GPUs it found and which backend it will use. Here is where to find it:

SystemWhere the log is
WindowsPress Windows + R, type explorer %LOCALAPPDATA%\Ollama, open server.log
Linux (Ubuntu, Kali, Fedora)journalctl -u ollama --no-pager
Maccat ~/.ollama/logs/server.log
Dockerdocker logs ollama (or your container's name)

Search the log for inference compute. That line names each device Ollama will use. When no GPU was found, it says the compute is the CPU. Here is how it reads on a machine where Ollama found no usable GPU, trimmed to the parts that matter:

msg="discovering available GPUs..."
msg="experimental Vulkan support disabled.  To enable, set OLLAMA_VULKAN=1"
msg="inference compute" id=cpu library=cpu compute="" name=cpu description=cpu

Three lines, one story. Ollama looked for GPUs, did not try Vulkan, and settled on library=cpu. On a machine where the GPU works, the same line names the card instead and shows library=CUDA, ROCm, Vulkan or Metal, along with how much memory the card has free.

That middle line deserves a second look, because it catches a lot of Intel and AMD owners. Older Ollama versions treated Vulkan as experimental and switched it off, so any GPU that only works through Vulkan was simply skipped. Current versions turn Vulkan on by default. If your log still says "experimental Vulkan support disabled," the fix is usually to update Ollama; if you must stay on that version, set OLLAMA_VULKAN=1 (how to set it is further down). Run ollama --version to see what you have.

Each server start also prints a server config line listing every Ollama setting. Glance through it for CUDA_VISIBLE_DEVICES, ROCR_VISIBLE_DEVICES or GGML_VK_VISIBLE_DEVICES. If any of them is set to -1, something on your system is deliberately hiding the GPU from Ollama, often a leftover from a tutorial. For more detail, set OLLAMA_DEBUG=1, restart Ollama, and the log will explain each GPU it rejected.

Cause 1: the model or its context doesn't fit in VRAM (the most common)

If ollama ps shows a split, Ollama is working exactly as designed: it filled your GPU and put the rest on the CPU. The fix is to make the job fit. What has to fit is more than the model file:

  • The model weights. Roughly the size you see in ollama list.
  • The context (KV cache). Memory for the conversation the model can see at once. It grows with the context length you set, and for long contexts it can be several gigabytes on its own.
  • A safety margin. Ollama keeps some VRAM free, and Windows, your browser and your display use some too.

Here is a rough guide for common sizes with the usual 4-bit download (the q4_K_M quantization that Ollama's default tags mostly use) and a modest context:

Your VRAMFits fully on the GPU, roughlyWill split onto the CPU
4 GB1B to 3B models7B and up
6 GBUp to about 7B8B with long context, 12B and up
8 GB7B to 8B comfortably12B to 14B
12 GBUp to about 12B, 14B with a short context14B with long context, 20B and up
16 GBUp to about 14B with room to spare27B to 32B
24 GBUp to about 32B70B

Treat the table as a starting point; every model and context differs, and ollama ps is the final word. Jake's case sat right on an edge: a 14B model on a 12 GB card would have fit with the default context, but the 32,000-token context he had copied from a video added several gigabytes more, and about a third of the model spilled to the CPU.

Ways to make it fit, cheapest first:

  1. Lower the context. Ollama's default is 4,096 tokens, which is plenty for chat. If you raised it, in Ollama's settings, with OLLAMA_CONTEXT_LENGTH, or with /set parameter num_ctx, bring it back down and check ollama ps again.
  2. Pick a smaller model or a smaller quantization. An 8B model fully on the GPU usually beats a 14B model half on the CPU, in speed and often in how it feels to use. Most models list smaller tags on their Ollama page.
  3. Close what else is using VRAM. Games, video editors, other AI apps, and browsers with hardware acceleration all hold graphics memory. nvidia-smi lists them.
  4. Unload other models. Ollama can keep several models loaded at once. ollama ps lists them all; ollama stop modelname unloads one.
  5. Turn on flash attention and a smaller KV cache. Setting OLLAMA_FLASH_ATTENTION=1 and OLLAMA_KV_CACHE_TYPE=q8_0 roughly halves the memory the context uses, with very little quality loss for most chat use. Restart Ollama after setting them.

🙋‍♂️ Jake's Reality Check

"But the video said a 32K context makes it smarter."

A bigger context lets the model see more text at once. It does not make each answer smarter, and for short customer replies you will never fill 4,096 tokens. Ethan set Jake's context back to the default; ollama ps switched to 100% GPU, and replies dropped from about 40 seconds to about 4.

Ollama CPU vs GPU: why even a small CPU share hurts so much

Jake's first question after the fix was a fair one: "If 69% was already on the GPU, why was it ten times slower, not just a bit slower?" The answer is memory speed, and it explains every split you will ever see in ollama ps.

To write each word, the model has to read through all of its weights once. So the speed limit is not how clever the chip is; it is how fast the chip can read its memory, called memory bandwidth. A mid-range card such as the RTX 3060 reads its VRAM at about 360 GB per second. A typical desktop or laptop reads system RAM at roughly 50 GB per second with dual-channel DDR4-3200, or about 90 GB per second with dual-channel DDR5-5600. That gap is the whole story.

A rough rule of thumb: the best possible speed in words per second is about the memory bandwidth divided by the size of the model being read. For a 5 GB model, that is a ceiling of around 70 per second from that card's VRAM and around 10 to 18 per second from system RAM. Real speeds land below those ceilings, but the ratio holds. When a model is split, every word has to wait for the CPU's slice too, read at RAM speed, so a third of the model on the CPU can cost far more than a third of the speed. That is why a smaller model fully on the GPU so often beats a bigger one that spills over.

The same rule tells you what to buy if you are choosing a card for Ollama: VRAM first, bandwidth second, everything else third. A used 12 GB card often runs local models better than a newer 8 GB one, because the model fits. Our guide to RAM and VRAM tiers for local AI goes through it by budget, and a mid-size model like Qwen3.8 is a good test of where your card's limit sits.

What is OLLAMA_GPU_OVERHEAD?

It is the amount of VRAM, in bytes, that Ollama leaves untouched on each GPU when it decides how many layers to load. The default is zero. Raising it makes Ollama more cautious, which helps if your display or another app needs headroom and models crash with out-of-memory errors. It never makes more of a model fit; it does the opposite. If your problem is a split you want to shrink, leave it at zero.

Ollama "not using full GPU" when VRAM looks free

Ollama estimates memory before it loads a model, and it estimates conservatively. If ollama ps shows a small CPU share even though nvidia-smi shows a gigabyte or two free, that margin is deliberate, to avoid crashing in the middle of a reply. Shortening the context by a few thousand tokens usually tips it to 100% GPU. You can also force the layer count with the num_gpu option (for example /set parameter num_gpu 99 inside ollama run to ask for every layer), but if the estimate was right, the model will fail to load or crash, so step back down if it does.

Cause 2: NVIDIA on Windows ("ollama not using nvidia gpu windows")

If ollama ps says 100% CPU and you have an NVIDIA card, check these in order.

  1. Driver version. Open a terminal and run nvidia-smi. The top line shows the driver version. Ollama needs 550 or newer, and 570 or newer for older cards. If nvidia-smi is not found at all, the NVIDIA driver is missing or broken. Install the current driver from NVIDIA's site or the NVIDIA App, restart, and quit and reopen Ollama.
  2. Card generation. Ollama needs CUDA compute capability 5.0 or newer. In plain terms: GTX 750 Ti and the GTX 900 series and newer work (GTX 900 series and GTX 10 series cards need the 570+ driver); older Kepler cards such as most of the GTX 600 and 700 series do not, and Ollama runs those PCs on the CPU.
  3. A hidden GPU. Check the server log's config line for CUDA_VISIBLE_DEVICES. If it is -1 or a number that does not match your card, remove it from Windows' environment variables and restart Ollama.
  4. Restart Ollama itself. Ollama detects GPUs when its server starts. If you installed or updated the driver while Ollama was running, right-click the Ollama icon in the system tray, choose Quit, and start it again; restarting the PC works too.
  5. Update Ollama. Old versions ship old CUDA libraries that newer cards, especially the RTX 50 series, need updates for. Download the current installer from ollama.com.
NVIDIA seriesCompute capabilityWorks with Ollama?
GTX 600 / most GTX 700 (Kepler)3.xNo, CPU only
GTX 750 Ti, GTX 900 series (Maxwell)5.0 to 5.2Yes, with driver 570 or newer
GTX 10 series (Pascal)6.1Yes, with driver 570 or newer
GTX 16 and RTX 20 series (Turing)7.5Yes, driver 550 or newer
RTX 30 series (Ampere)8.6Yes, driver 550 or newer
RTX 40 series (Ada)8.9Yes, driver 550 or newer
RTX 50 series (Blackwell)12.0Yes, with a current driver and a current Ollama

Laptops with two graphics chips

Many gaming and creator laptops have an Intel or AMD graphics chip plus an NVIDIA one. Ollama talks to the NVIDIA chip through CUDA directly, so Windows' per-app graphics preference does not decide it. What can stop it is a power mode that switches the NVIDIA chip off: some laptops have an "Eco," "Integrated only" or battery-saving GPU mode in the maker's control app (or a MUX switch setting). Plug in, choose the standard or "hybrid" mode, restart Ollama, and check nvidia-smi again. If nvidia-smi cannot see the card, neither can Ollama.

"Ollama not using GPU after update"

When it worked yesterday and not today, one of three things changed. A Windows Update may have swapped your NVIDIA driver for an older generic one; nvidia-smi will show a lower version or fail, and reinstalling the current driver fixes it. An Ollama update may need a newer driver than you have; the log's inference compute line will show the CPU, and updating the driver fixes that too. Or, rarely, the update changed which backend Ollama picks; the log names it, and the environment variables let you steer it. If a driver update itself broke your display, rolling it back safely takes two minutes.

Cause 3: NVIDIA on Linux ("ollama not using gpu ubuntu")

On Ubuntu, Kali, Fedora and other distributions, the same rule applies: if nvidia-smi does not work, Ollama cannot use the card. Linux adds three extra traps.

  • The open-source nouveau driver. Out of the box, many distributions use nouveau, which works for the desktop but not for CUDA. Install NVIDIA's proprietary driver. On Ubuntu, sudo ubuntu-drivers install picks the right one; on Kali, sudo apt install nvidia-driver nvidia-cuda-toolkit is the documented route; then reboot.
  • Secure Boot. If Secure Boot is on, the NVIDIA kernel module has to be signed. Ubuntu's installer asks you to set a password and enroll a key on the next reboot (the blue "MOK management" screen). Skip that screen and the module never loads, so nvidia-smi fails.
  • After sleep. On some systems the GPU disappears from Ollama after suspend and resume. Reload the NVIDIA memory module with sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm, then restart Ollama.

Ollama on Linux runs as a background service, so after any driver change, restart it with sudo systemctl restart ollama and check the log with journalctl -u ollama --no-pager | grep "inference compute". If you installed Ollama before the driver, the service may also have started before the GPU was ready; a restart fixes that too. On Kali in particular, if sudo itself keeps refusing you, our Kali basics guides start from the beginning.

Cause 4: AMD Radeon ("ollama not using amd gpu")

AMD is where the most outdated advice lives, because support has changed a lot. There are two routes into an AMD card, and which one you get depends on your card and your operating system.

ROCm is AMD's own compute platform, and Ollama uses it for cards on its supported list. That list is much shorter on Windows than on Linux:

SystemRadeon cards supported through ROCm
WindowsRX 7900 XTX, 7900 XT, 7900 GRE, 7800 XT, 7700 XT, 7600 XT, 7600, and Radeon PRO W7900 to W7500
LinuxFrom the RX 6800 up through the RX 9070 XT, plus Radeon PRO, Radeon AI PRO, Ryzen AI processors and Instinct accelerators

Vulkan is the second route, and it is the reason many "unsupported" Radeons now work. Current Ollama versions turn Vulkan on by default on Windows and Linux, so an RX 6600, an RX 5700 or a Radeon iGPU that ROCm does not list can still be used through Vulkan. If your log says "experimental Vulkan support disabled," update Ollama, or set OLLAMA_VULKAN=1 on an older version.

⚠️ The HSA_OVERRIDE_GFX_VERSION advice is Linux-only. Older guides tell every AMD user to set HSA_OVERRIDE_GFX_VERSION=10.3.0. On Linux, that trick makes ROCm accept some cards just outside its list, such as the RX 6600 and 6700 family, by telling ROCm to treat them as a close supported relative. It does nothing useful on Windows. On Windows, a card outside the ROCm list goes through Vulkan instead. If you set that variable on Windows, remove it.

Three more AMD checks on Linux:

  • Permissions. ROCm needs access to /dev/kfd and /dev/dri. The user running Ollama must be in the render and video groups. The standard install script handles this for the ollama service user; if you run ollama serve as yourself, add yourself with sudo usermod -aG render,video $USER and log in again.
  • Driver age. A timeout during GPU discovery in the log usually means the AMD driver is older than the ROCm version Ollama ships. Upgrade to the current ROCm-compatible driver.
  • Integrated graphics. Ryzen processors with Radeon graphics share system memory with the CPU. They can run small models through ROCm (on Ryzen AI chips) or Vulkan, but they will not be as fast as a separate card, and the "VRAM" they report is a slice of your RAM.

On Windows, the simplest health check is the same as for NVIDIA: keep the AMD Adrenalin driver current, quit and restart Ollama, and look for the inference compute line naming your Radeon. If an AMD driver update broke something else on the way, our display-driver guide covers rolling back cleanly.

Cause 5: Intel Arc and Iris Xe ("ollama not using intel gpu")

Intel GPUs work with Ollama through Vulkan. That covers Intel Arc desktop and laptop cards and the Arc and Iris Xe graphics built into recent Intel Core and Core Ultra processors. Three things decide whether it works:

  1. A current Ollama. In versions where Vulkan was still experimental, Intel GPUs were skipped unless you set OLLAMA_VULKAN=1. Update, or set it.
  2. A current Intel graphics driver. On Windows, get it from Intel's Driver & Support Assistant or your laptop maker. On Linux, keep the Mesa graphics stack current; Ubuntu and Kali ship it through normal updates.
  3. Linux permissions for memory reporting. If Ollama on Linux finds the Intel GPU but reports no usable memory, give the binary permission to read GPU memory counters: sudo setcap cap_perfmon+ep /usr/local/bin/ollama, then restart the service.

Set your expectations honestly: integrated Iris Xe and Arc graphics share system RAM, so they speed up small models but will not match a dedicated card with its own fast memory. A dedicated Arc card with 12 to 16 GB is a different story and handles 7B to 14B models well.

Ollama on a Mac

On an Apple Silicon Mac (M1 and every chip since), Ollama uses the GPU through Metal automatically, and there is nothing to configure. The GPU and CPU share the same unified memory, so the limit is your total memory, minus what macOS keeps for itself. If ollama ps shows a CPU share on a Mac, the model is simply too large for the memory available; pick a smaller tag. On older Intel-based Macs, Ollama runs on the CPU; their AMD and Intel graphics chips are not supported for this.

Ollama in Docker not using the GPU

A container cannot see your GPU unless you pass it in, and the way you do that differs by brand.

NVIDIA: install the NVIDIA Container Toolkit on the host first, then start Ollama with the --gpus=all flag:

docker run -d --gpus=all -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama

To check that Docker can reach the GPU at all, before blaming Ollama, run docker run --rm --gpus all ubuntu nvidia-smi. If that fails, the problem is the Container Toolkit or the host driver, not Ollama.

AMD: use the ROCm image and pass the two AMD devices:

docker run -d --device /dev/kfd --device /dev/dri -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama:rocm

On Fedora and other SELinux systems, the container may still be blocked from the devices; sudo setsebool container_use_devices=1 allows it. With Docker Compose, the same rules apply: a deploy.resources.reservations.devices entry for NVIDIA, or devices: entries for AMD. After the container starts, docker logs ollama shows the same inference compute line as a normal install.

Ollama in WSL2 on Windows

WSL2 runs Linux inside Windows, and NVIDIA GPUs work in it, with one rule people break constantly: install the NVIDIA driver in Windows only. Do not install a Linux NVIDIA driver inside WSL; the Windows driver provides the GPU to Linux automatically. In your WSL terminal, nvidia-smi should list the card. If it does, Ollama inside WSL can use it.

Two more things to know. WSL2 limits how much of your PC's RAM the Linux side can use (by default about half), which matters when a model spills to the CPU. And for most people who just want to chat with a model on Windows, the native Windows version of Ollama is simpler and avoids both issues. WSL2 makes sense when your other tools live in Linux. It also needs Windows' virtualization platform, which can clash with VirtualBox; our VT-x guide explains that trade-off.

How to set Ollama environment variables on Windows, Linux and Mac

Several fixes in this guide are environment variables. Ollama only reads them when its server starts, so after setting one, always restart Ollama.

Windows:

  1. Quit Ollama: right-click its icon in the system tray and choose Quit.
  2. Open Start, search for Edit environment variables for your account, and open it.
  3. Choose New, enter the name (for example OLLAMA_FLASH_ATTENTION) and value (1), and choose OK.
  4. Start Ollama again from the Start menu.

Linux (service install): run sudo systemctl edit ollama.service, add the setting under a [Service] line, save, then reload and restart:

[Service]
Environment="OLLAMA_FLASH_ATTENTION=1"
Environment="OLLAMA_KV_CACHE_TYPE=q8_0"

sudo systemctl daemon-reload
sudo systemctl restart ollama

Mac: run launchctl setenv OLLAMA_FLASH_ATTENTION 1 in Terminal, then quit and reopen the Ollama app.

The variables that matter most for GPU problems:

VariableWhat it does
OLLAMA_CONTEXT_LENGTHDefault context size for every model. Smaller uses less VRAM.
OLLAMA_FLASH_ATTENTION1 turns on flash attention, which saves memory for the context.
OLLAMA_KV_CACHE_TYPEq8_0 or q4_0 shrinks context memory (default f16).
OLLAMA_VULKAN1 enables Vulkan on older versions; 0 turns it off on current ones.
CUDA_VISIBLE_DEVICESWhich NVIDIA GPUs Ollama may use; -1 hides them all.
ROCR_VISIBLE_DEVICESWhich AMD GPUs Ollama may use through ROCm.
GGML_VK_VISIBLE_DEVICESWhich GPUs Ollama may use through Vulkan; -1 hides them.
HSA_OVERRIDE_GFX_VERSIONLinux only: lets ROCm accept a near-supported Radeon.
OLLAMA_MAX_LOADED_MODELSHow many models may stay loaded at once; lower it if models crowd each other out of VRAM.
OLLAMA_GPU_OVERHEADVRAM to keep free per GPU, in bytes. Raise it only to stop crashes.
OLLAMA_DEBUG1 makes the log explain GPU discovery in detail.

How to run Ollama on CPU only (on purpose)

Sometimes you want the CPU: the GPU is busy with a game or a render, it is too small for the model you need, or you are testing. Ollama does not need a GPU at all; on a CPU it is simply slower. To force it:

  • For one session: inside ollama run, type /set parameter num_gpu 0. Zero layers go to the GPU.
  • For one API call: add "options": {"num_gpu": 0} to the request.
  • For everything: hide the GPUs from the server and restart it, with CUDA_VISIBLE_DEVICES=-1 for NVIDIA, ROCR_VISIBLE_DEVICES=-1 for AMD through ROCm, and OLLAMA_VULKAN=0 or GGML_VK_VISIBLE_DEVICES=-1 for Vulkan devices.

Which models run acceptably on a CPU? Small ones. On a typical modern laptop CPU, 1B to 4B models answer at a readable pace, 7B to 8B models are usable if you are patient, and anything larger is a test of patience. Memory matters as much as cores: the whole model has to fit in RAM, with room for Windows. If you are choosing hardware for this, small models like Granite 4.2 show what a modest PC can do, and Ternary Bonsai shows how far compression now stretches a 6 GB card.

Two or more GPUs

With more than one GPU, Ollama loads a model onto a single card when it fits, because that is fastest. Only when it does not fit does it spread the model across cards. To always spread across all cards, set OLLAMA_SCHED_SPREAD=1. To use only some cards, list them in CUDA_VISIBLE_DEVICES (for example 0,1); NVIDIA recommends using the GPU's UUID, shown by nvidia-smi -L, because the numbering can change. If one card is much older or smaller than the other, it can slow the pair down; test with each card alone first.

Serving Ollama to a small team from one GPU

Jake eventually wanted his part-time technician to use the same model from the shop's second PC. Sharing one GPU box is common in small offices and labs, and it changes the memory math. Each parallel request needs its own context memory, so OLLAMA_NUM_PARALLEL multiplies the KV cache: four parallel users at 8,192 tokens use roughly the context memory of one user at 32,768. A model that ran at 100% GPU for one person can split onto the CPU as soon as the second person connects.

Practical rules for a shared box: keep the context modest, set parallelism to the number of people who really type at once, keep OLLAMA_MAX_LOADED_MODELS low so a second model does not push the first one off the GPU, and expose Ollama on the network only through a firewall rule for the machines that need it (OLLAMA_HOST=0.0.0.0 opens it to everyone on the network). When a team outgrows one card, renting a cloud GPU server by the hour is often cheaper than a second workstation, but the same VRAM rules apply there.

100% GPU but still slow? And a final checklist

If ollama ps says 100% GPU and replies still crawl, the GPU is doing its job and something else is the limit:

  • The first reply is always slower. Loading a model from disk into VRAM takes seconds to a minute; later replies are fast. Models unload after five idle minutes by default; OLLAMA_KEEP_ALIVE changes that.
  • A small or old card is still a small card. A GTX 1650 with 4 GB runs a 3B model on the GPU, but not quickly. That is normal.
  • Integrated graphics share RAM bandwidth with everything else, so they gain less than a dedicated card.
  • Laptops on battery often cap the GPU's power. Plug in and pick a performance power mode.

And if you are still on 100% CPU after reading this far, go through the list in order. It finds the cause almost every time:

  1. ollama ps with a model loaded: is it really 100% CPU, or a split?
  2. Server log: what does the inference compute line say?
  3. ollama --version: is it current? Does the log say Vulkan is disabled?
  4. Driver: does nvidia-smi (or the AMD/Intel tool) see the card, and is the driver current?
  5. Card: is it on the supported list for your operating system?
  6. Variables: is any *_VISIBLE_DEVICES set to -1, or a stray HSA_OVERRIDE_GFX_VERSION on Windows?
  7. Container or WSL: can nvidia-smi see the card inside it?
  8. Restart Ollama after every change, then check ollama ps again.

If a model refuses to load at all rather than loading on the CPU, that is a different problem, often a memory or architecture error; our guide to "failed to load model" decodes those messages, and most of it applies to Ollama too.

Ollama GPU: frequently asked questions

Does Ollama run on GPU or CPU?

Both. It uses a supported GPU automatically and falls back to the CPU for whatever does not fit in VRAM, or for everything if no usable GPU is found. Run ollama ps to see which one your model is using.

How do I check if Ollama is using my GPU?

Load a model, then run ollama ps and read the PROCESSOR column. 100% GPU means fully on the GPU, 100% CPU means not at all, and a split means part of the model did not fit.

Why is Ollama not using my GPU on Windows?

Usually the model or context is too big for your VRAM, the NVIDIA driver is older than 550, the card is too old or not on the AMD list, or an old Ollama skipped Vulkan. Check ollama ps, then the inference compute line in server.log.

Why does Task Manager show my GPU at 0% while Ollama runs?

Task Manager's default graphs show 3D work, not AI compute. Watch Dedicated GPU memory instead, or switch a graph to Cuda or Compute_0, or use nvidia-smi.

Does Ollama automatically use the GPU?

Yes. There is no setting to turn it on. If Ollama finds a supported GPU when its server starts, it loads models onto it first.

What are Ollama's GPU requirements?

NVIDIA cards with compute capability 5.0 or newer and driver 550 or newer (570 for older cards), AMD Radeon cards on its ROCm list or supported through Vulkan, Intel Arc and Iris Xe through Vulkan, and Apple Silicon Macs through Metal.

Why is Ollama not using my AMD GPU?

On Windows, ROCm supports only the RX 7000 series and Radeon PRO W7000 cards; other Radeons need Vulkan, which current Ollama versions enable by default. On Linux, check the ROCm list, the render and video groups, and the driver version.

Does HSA_OVERRIDE_GFX_VERSION work on Windows?

No. It is a Linux ROCm setting. On Windows, a Radeon outside the ROCm list is used through Vulkan instead, so remove the variable if you set it.

Why is Ollama not using my Intel GPU?

Intel GPUs work through Vulkan. Update Ollama, or set OLLAMA_VULKAN=1 on older versions, keep the Intel graphics driver current, and on Linux allow memory reporting with setcap cap_perfmon+ep on the ollama binary.

How do I make Ollama use the GPU in Docker?

For NVIDIA, install the NVIDIA Container Toolkit and start the container with --gpus=all. For AMD, use the ollama/ollama:rocm image with --device /dev/kfd and --device /dev/dri.

Why is Ollama only using part of my GPU?

The model plus its context did not fit in VRAM, so some layers run on the CPU. Lower the context, choose a smaller model, close other GPU apps, or enable flash attention with a q8_0 KV cache.

How do I run Ollama on CPU only?

Inside ollama run, type /set parameter num_gpu 0. To force it for everything, set CUDA_VISIBLE_DEVICES=-1 for NVIDIA, ROCR_VISIBLE_DEVICES=-1 for AMD, or OLLAMA_VULKAN=0, then restart Ollama.

Does Ollama need a GPU?

No. Ollama runs on a CPU alone; it is just slower. Small models of 1B to 4B parameters run acceptably on most modern CPUs.

What is OLLAMA_GPU_OVERHEAD?

It is how much VRAM, in bytes, Ollama keeps free on each GPU. The default is zero. Raising it prevents some out-of-memory crashes but means less of the model fits on the GPU.

What is the best Ollama model for an 8 GB GPU?

A 7B or 8B model in the default 4-bit download, with the default context, usually runs fully on an 8 GB card. Check with ollama ps; if it splits, try a smaller tag.

Why did Ollama stop using my GPU after an update?

A Windows Update may have replaced your graphics driver, or the new Ollama needs a newer driver. Check nvidia-smi and the inference compute line in the log, then update the driver and restart Ollama.

Can Ollama use two GPUs?

Yes. It uses one card when the model fits and spreads across cards when it does not. Set OLLAMA_SCHED_SPREAD=1 to always spread, or choose cards with CUDA_VISIBLE_DEVICES.

Does Ollama use the GPU on a Mac?

Yes, on Apple Silicon Macs it uses Metal automatically. Older Intel Macs run Ollama on the CPU.

Jake's card had been working from the first day; it was simply being asked to hold more than it could. With the context back at the default, his 14B model sits at 100% GPU, quotes come back in a few seconds, and nothing about his customers leaves the shop. If you have spent an evening reinstalling drivers over this, you did nothing wrong. Ollama's quiet fallback to the CPU looks exactly like a broken GPU until you know which column to read.

📌 If you keep one line from this page

Don't trust the graph; trust ollama ps.

A split means it didn't fit; 100% CPU means it never found the GPU.

Revision note. Written October 5, 2026. If your new graphics card seemed to sit idle, it was probably working all along; may your next reply arrive in seconds.

Related