Ollama "Model Requires More System Memory" and "Llama Runner Process Has Terminated": Every Fix
"Error: model requires more system memory (X GiB) than is available (Y GiB)" means Ollama estimated how much memory the model needs, compared it with the memory that is free right now, and refused to load it. Here is the part that surprises people: your PC did not run out of memory. Ollama checked first and decided not to try. And on Windows with WSL2 or Docker, the "available" number can be about half your RAM, because WSL2 limits itself to 50% of your memory by default. Its close cousin, "llama runner process has terminated," is different: Ollama did try, and the worker process that holds the model crashed, most often because it ran out of RAM or graphics memory. The fixes overlap: free memory, pick a smaller model or a shorter context, unload models you are not using with ollama stop, make sure your GPU is being used, and raise the WSL2 or Docker memory limit if you run Ollama inside one. Below, we start with the basics and then fix every memory error Ollama shows, plus how to keep a model loaded or unload it on purpose.
Jake's shop setup had grown. After the connection fixes, his Ollama and Open WebUI ran together in Docker Desktop on the shop's Windows PC, which has 16 GB of RAM and a 12 GB graphics card. His technician wanted a bigger, smarter model for writing repair explanations, downloaded one that was about 9 GB, and got a red box in the chat page: "500: model requires more system memory (9.6 GiB) than is available (6.3 GiB)." Task Manager said Windows had more than 8 GB free. Jake closed Chrome, tried again, and got the same message with slightly different numbers. The technician went back to writing explanations by hand for the afternoon. When Ethan looked that evening, the 16 GB PC turned out to be giving Docker only 8 GB, and the GPU was not being used inside the container at all. Two settings, one restart, and the model loaded.
The basics: RAM, VRAM, and what "loading a model" uses
Every Ollama memory error is about one of two pools of memory, and knowing which one is in trouble is half the fix.
- RAM (system memory) is your computer's main working memory, the number you see as "16 GB" or "32 GB" in your PC's specs. The CPU uses it, and so does every app you have open.
- VRAM (graphics memory) is separate memory on a graphics card. The GPU uses it, and only what fits there runs at full GPU speed. Laptops without a separate card, and Apple Silicon Macs, share one pool between the two.
- Swap, or the page file, is disk space the system uses as emergency overflow when RAM is full. It keeps programs alive, but it is many times slower than real memory, so a model running from it crawls.
When Ollama loads a model, it needs room for three things at once: the model's weights (roughly the file size shown by ollama list), the context (memory for the conversation the model can see, which grows with the context length), and some working space for the calculations. Whatever does not fit in VRAM goes to RAM. If both are too small, you get one of the errors in this guide.
"So a 9 GB model needs more than 9 GB," Jake said.
"Always," Ethan said. "Think of moving into a flat. The furniture is the model. The context is the boxes you have not unpacked yet. And you still need floor space to walk around. Measure the flat for all three, not just the sofa."
The runner: why the error says "llama runner"
Ollama is not one program doing everything. The Ollama server you start stays small; for each model it loads, it launches a separate worker process, called a runner, which holds the model in memory and does the work. That design is why "llama runner process has terminated" is survivable: the worker crashed, but the server is still running and can try again. It is also why the useful details are never in the error itself but in the server log, a few lines above, where the runner said what went wrong before it died.
"model requires more system memory than is available": what it means and the fixes
This message comes from Ollama's safety check. Before loading, it estimates the model's needs and compares them with the memory that is free at that moment. The two numbers in the message are those two sides:
Error: model requires more system memory (9.6 GiB) than is available (6.3 GiB)
- The first number is what Ollama thinks the part of the model going into RAM will need, including its context and working space.
- The second number is RAM that is free right now. Not your total RAM, and not what Windows could free by closing things; what is free this second.
From Open WebUI and other apps, the same check arrives as an HTTP error, so it often shows up as "500: model requires more system memory..." The fix is the same. Work through these in order; each one either shrinks the first number or grows the second.
- Check that your GPU is actually being used. If Ollama found no usable GPU, the entire model has to fit in RAM, which is why a model that "should" fit on a 12 GB card is rejected. Run
ollama psafter loading any small model; if it says 100% CPU on a PC with a graphics card, fix that first with our GPU guide. This was half of Jake's problem: Docker had been started without access to the GPU. - Unload models you are not using. Ollama can keep several models loaded at once, and each one holds memory.
ollama pslists them;ollama stop modelnameunloads one. SettingOLLAMA_MAX_LOADED_MODELS=1keeps it to one at a time. - Free RAM. Close browsers with many tabs, games, video editors and other AI tools, then try again. On Windows, Task Manager's Memory page shows "Available"; on Linux,
free -hshows the available column, which is the one that matters. - Shorten the context. If you or an app raised the context length (with
OLLAMA_CONTEXT_LENGTH,num_ctxor a setting in Open WebUI), bring it back toward the 4,096-token default. Long contexts can add several gigabytes. - Choose a smaller model or a smaller quantization. A 14B model needs roughly twice the memory of a 7B one. Most models offer smaller tags on their Ollama page, and a 4-bit download (the usual default) needs far less than an 8-bit or 16-bit one.
- Raise the hidden limit, if you run Ollama in WSL2 or Docker. This is the other half of Jake's problem, and it gets its own section because it fools so many people.
- Add memory, as a last resort. On Linux, extra swap can let a borderline model load (it will be slow). The honest long-term fix is more RAM, or a model that suits the RAM you have.
Apps can raise the context behind your back
A frequent reason for "it worked in the terminal but fails in the app" is that the app asks for a much bigger context than the terminal does. Ollama's own default is 4,096 tokens, but Open WebUI, coding assistants and agent tools often send their own num_ctx with every request, sometimes 32,000 tokens or more, because longer documents and codebases need it. Each request with a different context can also make Ollama reload the model with the new size, so the same model briefly needs memory twice. In Open WebUI, look for a Context Length (num_ctx) value in the model's advanced parameters and in the chat's Controls panel; in coding tools, look for a context or "max tokens" setting in the model configuration. Lower it to what you really need, and the memory error often disappears without changing anything in Ollama.
The server log shows what was actually requested: each model load prints the context size it was given. If the terminal loads a model with 4,096 and the app with 131,072, you have found your culprit. Jake's Open WebUI had a 32,000-token context set on the technician's model from an earlier experiment, which on its own added several gigabytes to every load.
"evicting a model to make space" in the log
If the server log says "model requires more system memory than is currently available, evicting a model to make space," nothing is wrong. Ollama found another model loaded, unloaded it to make room, and carried on. You only see the error when there was nothing left to evict and the new model still did not fit.
How much RAM does Ollama need?
Ollama itself needs very little; the model decides. A practical rule for the usual 4-bit downloads: the model needs about its file size, plus 1 to 3 GB for context and working space, and your operating system and apps need their own share on top. Here is what that means on common PCs when the model runs on the CPU (no usable GPU):
| Your RAM | Comfortable on the CPU, roughly | Expect "requires more system memory" |
|---|---|---|
| 8 GB | 1B to 3B models | 7B and up, especially with apps open |
| 16 GB | Up to 7B to 8B | 14B with long context, 20B and up |
| 32 GB | Up to about 14B, 27B with care | 32B with long context, 70B |
| 64 GB | Up to about 32B, a 70B 4-bit model tightly | Anything much larger |
With a working GPU, the picture improves: VRAM holds as much of the model as fits, and only the rest needs RAM. A PC with 16 GB of RAM and a 12 GB card can run models that a 32 GB PC without a GPU struggles with. Treat the table as a starting point; ollama ps after loading tells you the real size, in its SIZE column. For a fuller picture by budget, see our RAM and VRAM tiers for local AI.
Why Ollama seems to use too little, or too much, memory
Two opposite puzzles come up constantly. The first: "I have 64 GB of RAM, so why does Ollama only use 6?" Usually because the model is on the graphics card, where Task Manager and top count it as GPU memory, not RAM; ollama ps will say 100% GPU. The other reason is how the model file is read. Ollama can map the model file into memory rather than copying it into its own space, so system monitors may count those gigabytes as file cache instead of as Ollama's memory. On Linux, that shows up under buff/cache in free -h. The memory is in use; it is just filed under a different heading.
The second puzzle: "Ollama is using far more than the model's size." That is the context and the working space, multiplied by every model that is still loaded and every parallel request slot. A 5 GB model with a 32,000-token context, loaded twice by two apps that ask for different settings, can easily hold 15 GB. ollama ps lists every loaded model with its real size, and that is the number to trust. More RAM does not make a model smarter or faster once it fits; it only lets you load bigger models, or more of them at once.
How much storage does Ollama use?
Disk space is a separate question from memory. Each model you download takes roughly its listed size on disk, and they add up quickly: five mid-size models can be 30 to 50 GB. ollama list shows each model's size, and ollama rm modelname deletes one from disk. The models live in %HOMEPATH%\.ollama\models on Windows, ~/.ollama/models on a Mac, and /usr/share/ollama/.ollama/models for the Linux service. To store them on a bigger drive, set OLLAMA_MODELS to a folder there and restart Ollama.
The hidden limit: Ollama in WSL2 or Docker sees only part of your RAM
If Ollama runs inside WSL2, or inside Docker Desktop on Windows (which uses WSL2 underneath), it does not see your PC's memory. It sees the memory of a small virtual machine, and by default WSL2 gives that machine 50% of your total RAM. On a 16 GB PC, that is 8 GB, minus what Linux and Docker themselves use. That is exactly how Jake's PC showed more than 8 GB free in Task Manager while Ollama, inside Docker, saw about 6.
To raise it, open the WSL Settings app from the Start menu and increase the memory size, or create a file named .wslconfig in your user folder (%UserProfile%) with:
[wsl2] memory=12GB swap=8GB
Then run wsl --shutdown in PowerShell, wait about eight seconds, and start WSL or Docker Desktop again. The change only applies after that full shutdown. Leave Windows enough to breathe: on a 16 GB PC, 10 to 12 GB for WSL is the sensible ceiling. The swap line sets WSL's own overflow file (25% of your RAM by default); a little more helps borderline models load, at the cost of speed.
Docker containers can be capped too. On Docker Desktop for Mac, or with the Hyper-V backend on Windows, the limit is under Settings → Resources → Memory. On any system, a container started with --memory, or a Compose file with a memory limit, gets only that much, and Ollama's check reports that number. If you did not mean to limit it, remove the flag.
And do not forget the GPU when moving Ollama into Docker. A container started without --gpus=all (NVIDIA) or the ROCm devices (AMD) has no GPU at all, so every model lands in the container's RAM, the smallest pool on the machine. For Jake, Ethan raised WSL's memory to 12 GB and added the GPU to the Compose file; ollama ps then showed the 9 GB model at 100% GPU, and system RAM barely moved.
"llama runner process has terminated": find the real reason
This error means the model's worker process crashed during loading or while answering. The words after the colon are a clue, and the server log has the full story. Find the log first:
- Windows: press Windows + R, run
explorer %LOCALAPPDATA%\Ollama, openserver.log. - Linux:
journalctl -u ollama --no-pager -n 100 - Mac:
cat ~/.ollama/logs/server.log - Docker:
docker logs ollama
Scroll to the moment of the crash and read the ten or twenty lines above "llama runner process has terminated." That is where the runner explained itself. Then match the ending of the error:
| The error ends with | What usually happened | Fix |
|---|---|---|
| signal: killed | Linux ran out of RAM and its out-of-memory killer stopped the runner. | The memory fixes above; smaller model or context |
| CUDA error: out of memory, or cudaMalloc failed: out of memory | The graphics card ran out of VRAM while loading or answering. | VRAM fixes |
| exit status 2 | The runner stopped on an error of its own. The cause is in the log lines above: often an unsupported model, a damaged download, or memory. | The non-memory causes, below |
| signal: aborted (core dumped) | The engine hit a check it could not pass and stopped itself; often memory or an unsupported model. | Read the line above; update Ollama |
| exit status 0xc0000409 (Windows) | Windows' way of reporting that the runner aborted itself; same causes as the line above. | Read the log; update Ollama and the GPU driver |
| exit status 0xc000001d (Windows) | "Illegal instruction": the CPU lacks an instruction the code expected. Rare, usually on very old CPUs. | Update Ollama; check your CPU's age |
If the log is too quiet to tell, quit Ollama, set OLLAMA_DEBUG=1, start it again, and reproduce the crash. The debug log explains each step of loading, including how much memory it planned for each device.
"llama runner process has terminated: signal: killed" on Linux
This one is almost always memory. Confirm it in the kernel log, which records every process the out-of-memory killer stops:
journalctl -k | grep -iE "out of memory|killed process" # or sudo dmesg | grep -iE "out of memory|killed process"
A line naming the Ollama runner is your confirmation. Then use the fixes from the system-memory section: free RAM, unload other models, shorten the context, choose a smaller model, or add swap. To add an 8 GB swap file on Ubuntu, Kali or Debian (it will make borderline loads possible, not fast):
sudo fallocate -l 8G /swapfile sudo chmod 600 /swapfile sudo mkswap /swapfile sudo swapon /swapfile free -h # the Swap line should now show 8.0Gi # to keep it after reboot, add this line to /etc/fstab: # /swapfile none swap sw 0 0
exit status 2 and other crashes: the non-memory causes
When the log above the crash does not mention memory, three other causes cover most of the rest:
- The model is newer than your Ollama. A model released last week may use an architecture your installed version does not know yet. The log says something like unknown model architecture or fails while reading the model's metadata. Update Ollama; the download page always has the current version.
- The download is damaged. An interrupted pull or a full disk can leave a broken model file. Delete it and pull it again:
ollama rm modelname, thenollama pull modelname. - A GPU driver problem. Crashes that start right after a driver update, or only when the GPU is used, point at the driver. Update it (or roll it back), restart, and test once with the GPU hidden (
CUDA_VISIBLE_DEVICES=-1) to confirm the model runs on the CPU.
If you imported a model yourself from a GGUF file with a Modelfile, check that the file is complete and that the architecture is one Ollama supports; a model that fails in Ollama but loads elsewhere is usually ahead of Ollama's support. Our LM Studio load-error guide walks through reading those architecture errors, and most of it applies here.
CUDA "out of memory" and "memory layout cannot be allocated": graphics memory errors
These come from the graphics card side. Ollama planned how much of the model to put in VRAM, and at load time, or partway through a long answer, the card ran out. Common reasons: another app grabbed VRAM after Ollama measured it, the context grew past the estimate, or several requests ran in parallel and each needed its own context memory.
Fixes, cheapest first:
- Close other GPU apps. Games, video editors, other AI tools and browsers with hardware acceleration all hold VRAM.
nvidia-smilists what is using the card. - Shorten the context and keep
OLLAMA_NUM_PARALLELat 1 unless several people really use it at once; each parallel slot reserves its own context memory. - Enable flash attention and a smaller context cache:
OLLAMA_FLASH_ATTENTION=1andOLLAMA_KV_CACHE_TYPE=q8_0, then restart Ollama. Together they roughly halve the context's memory. - Leave headroom.
OLLAMA_GPU_OVERHEADtells Ollama to keep some VRAM free on each card. It takes bytes:1073741824is 1 GiB. A little headroom stops crashes mid-answer at the cost of putting slightly less of the model on the GPU. - Remove forced layer counts. See the next heading.
"memory layout cannot be allocated with num_gpu = N"
This message means Ollama was told to put a fixed number of layers on the GPU (that is what num_gpu means here) and the plan does not fit. It almost always comes from a setting someone added on purpose: a PARAMETER num_gpu line in a Modelfile, a num_gpu option sent by an app, or a value set with /set parameter. Remove it and let Ollama work out the split itself; it is good at it. If you set it to squeeze more speed out of the card, lower the number until the model loads, or shorten the context so the layers fit.
"timed out waiting for llama runner to start"
This one is about time, not size. Ollama gives a model five minutes to load by default. A large model on a slow hard drive, a network drive, or a PC that is already swapping can take longer, and the load is abandoned with the progress percentage it reached. Fixes: store models on an SSD (move them with OLLAMA_MODELS), check that antivirus software is not scanning the multi-gigabyte model files as they are read, free RAM so the system is not swapping, or raise the limit with OLLAMA_LOAD_TIMEOUT=15m. If the percentage barely moves, the PC is swapping; a smaller model is the real fix.
Keep a model in memory, or unload it from memory
Two of the most searched Ollama questions are opposites, and they use the same setting. By default, Ollama keeps a model loaded for five minutes after its last use, then unloads it to free memory. That is why the first answer after a coffee break is slow: the model is being loaded again.
To keep a model in memory longer:
- For everything: set
OLLAMA_KEEP_ALIVE=1h(or24h, or-1for forever) on the server and restart Ollama. - For one session:
ollama run modelname --keepalive 1h. - From an app or script: send
"keep_alive": "1h", or-1for forever, in the API request.
To unload a model from memory now:
- Run
ollama stop modelname. It is gone from RAM and VRAM immediately, and still on disk. - From the API, send a request for that model with
"keep_alive": 0. - Restarting Ollama unloads everything.
⚠️ "Remove from memory" is not "remove from disk." ollama stop unloads a model from memory and keeps the file. ollama rm deletes the model from disk, and you would have to download it again. If your goal is to free RAM for another model, you want ollama stop.
Keeping models loaded forever has a cost: they hold their memory even while idle, which can cause "requires more system memory" for the next model you load. On a shared PC, an hour is a good middle ground. For Jake's shop, Ethan set the keep-alive to two hours during opening hours, so the first quote of the morning is the only slow one.
Does Ollama use "shared GPU memory"?
Task Manager on Windows shows two numbers for a graphics card: Dedicated GPU memory (the card's own VRAM) and Shared GPU memory (a slice of system RAM that Windows lets the GPU borrow, usually up to half your RAM). People often hope Ollama will spill into the shared part when VRAM runs out. With a separate NVIDIA or AMD card, it does not, on purpose: layers that do not fit in VRAM run on the CPU from RAM instead, which is faster than making the GPU reach across into system memory for every word.
One related NVIDIA setting is worth knowing. Recent NVIDIA drivers on Windows include a CUDA - Sysmem Fallback Policy option in the NVIDIA Control Panel, under Manage 3D settings. It controls whether CUDA programs may spill into system memory when VRAM is full. If another AI app on your PC becomes suddenly, painfully slow rather than crashing when it runs out of VRAM, that fallback is usually why, and setting it to Prefer No Sysmem Fallback makes it fail clearly instead.
Integrated graphics are the exception to all of this. On laptops without a separate card, and on Apple Silicon Macs, "GPU memory" is shared system memory; there is only one pool. On AMD laptops and mini PCs with Radeon graphics built in, some BIOS screens offer a UMA frame buffer size setting that reserves more RAM for the graphics side, which can help Ollama place more of a model on the integrated GPU.
On a Mac: the unified memory ceiling
Apple Silicon Macs share one pool of memory between the CPU and the GPU, so there is no separate VRAM to run out of. There is, however, a ceiling: macOS lets the GPU use only part of the total, roughly two-thirds to three-quarters depending on how much memory the Mac has, and keeps the rest for the system. A 32 GB Mac therefore cannot give a 30 GB model to the GPU, and Ollama plans around that limit. Advanced users sometimes raise it with sudo sysctl iogpu.wired_limit_mb= followed by a number of megabytes; the change resets at the next restart, and setting it too high can make macOS itself unstable, so leave several gigabytes for the system. For most people, the better fix is the same as on a PC: a smaller model or a shorter context.
Thinking models fill the context faster
Many recent models "think" before they answer, writing out a long chain of reasoning first. That reasoning lives in the context too, so a thinking model can hit memory and speed limits on questions that a regular model answers comfortably. If you do not need the reasoning, turn it off: ollama run modelname --think=false, type /set nothink inside a session, or send "think": false in an API request, for models that support the switch. Answers come back shorter and faster, and the context stays small.
Ollama models that run well on a CPU only
If your PC has no usable GPU, or not much RAM, the fix for every error above is a model sized for it. Good choices for CPU-only machines are the small, recent models built for laptops and phones: 1B to 4B models such as Llama 3.2 in its 1B and 3B sizes, the small Gemma and Granite models, and other "mini" or "small" tags on Ollama's library. They answer at a readable pace on most modern CPUs and leave room for your other apps.
Three settings make any model friendlier to a CPU: a short context (the default 4,096 tokens is fine for chat), the default 4-bit download rather than 8-bit or full-size versions, and keeping only one model loaded. If you want to see how far a modest machine can go, our Granite 4.2 guide runs a capable model on an ordinary laptop, and Ternary Bonsai shows how compression squeezes a 27B model into a small memory budget.
Windows itself is low on memory
Sometimes the Ollama error is just the first symptom of a PC that is short on memory overall: Windows shows its own "Your computer is low on memory" warning, apps slow down, and the page file grows. Ollama's check counts on memory being genuinely free, so a PC with many background apps rejects models it could otherwise run. Our guide to the Windows low-memory warning covers finding what is using RAM, the page file settings, and when more RAM is the real answer.
How to uninstall Ollama and free the disk space
If you are done with Ollama, or want a truly clean reinstall, remove both the program and the models, which are stored separately and usually take far more space.
Windows: quit Ollama from the tray, then open Settings → Apps → Installed apps, find Ollama and choose Uninstall. The models stay behind in %HOMEPATH%\.ollama; delete that folder to get the space back, and delete %LOCALAPPDATA%\Ollama to remove the logs.
Linux (installed with the official script):
sudo systemctl stop ollama sudo systemctl disable ollama sudo rm /etc/systemd/system/ollama.service sudo rm $(which ollama) sudo rm -r /usr/local/lib/ollama # GPU libraries, if present sudo rm -r /usr/share/ollama # the service's models sudo userdel ollama sudo groupdel ollama
If you also ran ollama serve as yourself at some point, a second set of models may sit in ~/.ollama; check its size with du -sh ~/.ollama before deleting it.
Mac: quit Ollama from the menu bar, drag Ollama from Applications to the Trash, and delete the ~/.ollama folder for the models. If you installed it with Homebrew instead, run brew services stop ollama and brew uninstall ollama.
To keep Ollama but reclaim space, you do not need any of that: ollama list, then ollama rm the models you no longer use.
For IT admins: sizing a shared Ollama box
On a server that several people use, memory errors usually come from the multiplication nobody planned for. Each loaded model holds its weights; each parallel request slot holds its own context; and each additional model loaded "just for a minute" takes its share until its keep-alive expires. Size for the worst hour, not the average one: the largest model, times the context you allow, times OLLAMA_NUM_PARALLEL, plus headroom for the operating system.
Practical settings for a team box: pin OLLAMA_MAX_LOADED_MODELS to what memory can really hold, set a modest default context with OLLAMA_CONTEXT_LENGTH and let power users raise it per request, keep a sensible OLLAMA_KEEP_ALIVE, and watch memory with your normal monitoring. When the team outgrows one machine, adding a second GPU server, or renting cloud GPU instances by the hour for heavy jobs, is usually cheaper than buying the biggest card on the market, and the same memory math applies in the cloud.
Still failing? The memory checklist
- Read the exact error: "requires more system memory" (refused before loading) or "runner process has terminated" (crashed after trying)?
- Run
ollama pswith a small model: is your GPU being used at all? - Unload other models with
ollama stop, and close memory-hungry apps. - Check for a hidden limit: WSL2 (50% of RAM by default) or a Docker memory cap.
- Shorten the context and remove any forced
num_gpu. - Try a smaller model or quantization of the same family.
- Read the server log above the crash; on Linux, check the kernel log for the out-of-memory killer.
- Update Ollama and your GPU driver, then re-pull the model if it still crashes.
Ollama memory errors: frequently asked questions
What does "model requires more system memory than is available" mean in Ollama?
Ollama estimated the model needs more RAM than is free right now and refused to load it. Free memory, unload other models, shorten the context, choose a smaller model, or raise the WSL2 or Docker memory limit.
Why does Ollama say not enough memory when I have enough RAM?
It counts only free memory, not total RAM. Inside WSL2 or Docker Desktop on Windows, it sees only the virtual machine's share, which is 50% of your RAM by default. If no GPU is used, the whole model must fit in RAM.
What does "500: model requires more system memory" mean?
It is the same memory check, delivered as an HTTP error to an app such as Open WebUI. The fixes are the same as for the command-line error.
What does "llama runner process has terminated" mean?
The worker process that holds the model crashed while loading or answering. The words after the colon, and the server log lines just above, say why; memory is the most common cause.
What does "llama runner process has terminated: exit status 2" mean?
The runner stopped on an error of its own. Read the log lines above it: often the model is newer than your Ollama, the download is damaged, or memory ran out. Update Ollama or re-pull the model.
What does "signal: killed" mean in Ollama?
On Linux, the system ran out of RAM and the out-of-memory killer stopped the runner. Confirm with journalctl -k, then free memory, use a smaller model or context, or add swap.
How do I fix CUDA error out of memory in Ollama?
Close other GPU apps, shorten the context, keep OLLAMA_NUM_PARALLEL at 1, enable flash attention with a q8_0 KV cache, and set OLLAMA_GPU_OVERHEAD to leave some VRAM free.
What does "memory layout cannot be allocated" mean?
A fixed number of GPU layers was requested with num_gpu and does not fit in VRAM. Remove the num_gpu setting from your Modelfile, app or session and let Ollama choose.
How much RAM does Ollama need?
About the model's file size plus 1 to 3 GB, on top of what your system uses. 8 GB suits 1B to 3B models, 16 GB suits 7B to 8B, and 32 GB suits up to about 14B on the CPU.
How do I keep an Ollama model in memory?
Set OLLAMA_KEEP_ALIVE to a duration such as 1h, or -1 for forever, and restart Ollama. For one session, use ollama run modelname --keepalive 1h, or send keep_alive in the API request.
How do I unload a model from memory in Ollama?
Run ollama stop modelname. It frees RAM and VRAM immediately and keeps the model on disk. From the API, send a request with keep_alive set to 0.
How do I remove a model from Ollama completely?
Run ollama rm modelname. That deletes it from disk; you would need to pull it again. Use ollama stop instead if you only want to free memory.
Does Ollama use shared GPU memory?
Not with a separate graphics card: layers that do not fit in VRAM run on the CPU instead. On integrated graphics and Apple Silicon, system memory is the GPU's memory, so it is used directly.
How do I give WSL2 more memory for Ollama?
Use the WSL Settings app, or add memory=12GB under [wsl2] in %UserProfile%\.wslconfig, then run wsl --shutdown and restart WSL or Docker Desktop.
Why does Ollama say "timed out waiting for llama runner to start"?
The model took longer than the five-minute load limit, usually because of a slow disk, antivirus scanning, or a PC that is swapping. Use an SSD, free memory, or raise OLLAMA_LOAD_TIMEOUT.
How much storage does Ollama use?
Each model takes about its listed size on disk. ollama list shows sizes, ollama rm deletes models, and OLLAMA_MODELS moves storage to another drive.
Which Ollama models can run on a CPU only?
Small models of 1B to 4B parameters, such as Llama 3.2 1B and 3B and the small Gemma and Granite models, run well on most modern CPUs with the default 4-bit download and a short context.
Does adding swap fix Ollama memory errors?
On Linux it can let a borderline model load, but the model will be very slow while it uses swap. A smaller model or more RAM is the real fix.
Jake's 16 GB PC was never too small for that model. Docker was quietly working inside an 8 GB box, and the graphics card was sitting outside it. With the WSL limit raised, the GPU passed into the container and a two-hour keep-alive, the technician's bigger model loads once in the morning and answers in seconds all day. If a memory error has made your PC feel underpowered, take heart: most of the time, the memory was there all along, just behind a limit nobody mentioned.
📌 If you keep one line from this page
"Requires more memory" means Ollama checked first; "runner terminated" means it tried and crashed.
In WSL2 and Docker, check the hidden limit before blaming your RAM.
Revision note. Written October 5, 2026. If a memory error made your PC feel too small, it may only have been fenced in; may your next model load on the first try.