THINK BIG
Local AI Deployment
Is your local AI actually using the GPU? "Installed" doesn't mean "running well"
When a local AI model is painfully slow — minutes per response, or constant timeouts — the most likely cause isn't underpowered hardware. It's that the model is running entirely on the CPU, and the GPU is sitting idle. One command, ten seconds, and you'll know for certain.
(Note: This article uses Ollama as the illustrative example because it is one of the most common tools for running local models on a personal computer. The underlying principles apply to other runtimes; the specific commands differ.)
1. The Situation: A well-documented gap
Running an AI model on your own machine — no subscription, no data leaving the building, available offline — has gone from an enthusiast experiment to a realistic option for individuals and small teams. The most common complaint that followed is a single sentence: "I installed it, but it's too slow to use."
The open-source community has documented this thoroughly. Ollama's GitHub issue tracker includes reports of models that "occasionally fall back to CPU, causing 100% CPU utilization, until a restart." (ollama/ollama #13765) Another thread describes a version update that caused models to stop using the GPU entirely. (ollama/ollama #13320) Published troubleshooting guides note that this failure mode is particularly insidious because it produces no error message — an outdated driver, a disabled environment variable, and the model quietly falls back to CPU, while everything on screen looks normal. (cmaven.github.io, localaimaster.com)
In other words: a "slow local AI" is usually not a performance problem. It's a configuration problem that hasn't been noticed yet.
2. The Debate: Is local deployment worth it?
The case for local deployment is concrete: data stays on your hardware, there's no per-token bill anxiety, and the model can run around the clock as part of an AI Agent workflow. For anyone who cares about privacy or wants continuous AI presence without ongoing cloud costs, this is something cloud services simply can't offer.
The reservations cluster around two points:
The setup threshold is higher than expected. Driver version, environment variables, model size relative to video RAM, memory allocation — any one of these being wrong degrades the result significantly, and there's no obvious starting point for diagnosis.
Hardware has to match the model. If the model is larger than the available video RAM, the overflow gets offloaded to the CPU. The official Ollama FAQ notes that models can end up in a "partial CPU, partial GPU" state. (docs.ollama.com/faq) That split slows everything down.
Both sides are pointing at the same thing: local AI can be excellent, but only after someone confirms it's actually running on the right hardware.
3. What we learned from actually doing this
"Looks fine" is the most dangerous state. We had a machine where the model was installed, conversations worked, and nothing reported an error — but every response took one to four minutes, with occasional timeouts. We spent considerable time adjusting other settings before finally checking: the model had been running 100% on the CPU. The GPU was untouched. We now check GPU utilization before touching anything else.
One metric is all you need. ollama ps shows the PROCESSOR column — 100% GPU, 100% CPU, or a split. The API also returns video RAM usage (size_vram). If size_vram is zero, the model is on CPU, regardless of what everything else looks like. We screenshot this number before every delivery.
Three root causes we've actually encountered — none requiring a hardware upgrade:
The GPU wasn't detected. Some newer integrated or hybrid graphics architectures aren't automatically recognized. They require additional environment configuration, and that configuration must be set for the specific user account that runs the model, not just the admin account.
The install path contained non-ASCII characters. The detection process crashed silently when the path included characters outside standard ASCII. Moving to a plain ASCII path fixed it immediately — the kind of trap that never appears in any spec sheet.
The context window was set too large. Setting the AI's memory (context window) too high exhausted shared system RAM and caused cascading system instability. Ollama defaults to 4096 tokens (per the official FAQ) — that's a reasonable baseline. Increasing it is fine, but it should be sized to the machine.
The difference after fixing it. Same machine, same model — switching from CPU-only to actual GPU use took response time from "minutes per question" to "output starts within seconds." That's a configuration fix, not a hardware upgrade.
4. Our recommendation: Three steps when local AI is slow
Check the PROCESSOR column first — before anything else. Run ollama ps. 100% CPU means the GPU isn't being used; a CPU/GPU split usually means the model is larger than available video RAM. This takes ten seconds and can save hours.
Match model size to video RAM. Bigger isn't always better. A smaller model that fits entirely in video RAM will often outperform a larger model that spills onto the CPU. Choose based on what your hardware can actually hold.
Make "confirm the GPU" a standing check. Driver updates, software updates, and account changes can all silently revert to CPU. Check after any meaningful change — don't wait for "why is this slow again?" to surface the issue.
One line to close with: when local AI is slow, suspect the configuration before the hardware. In most cases, the machine is capable enough — it just hasn't been used correctly yet.
What "delivered and running well" actually means
When we deploy a local AI Agent on a client's machine, "GPU actually running" is a mandatory pre-delivery check — not something we do only when a problem surfaces. Our delivery standard includes:
Evidence before handoff, not a verbal "it's installed." We confirm which hardware the model is using and record the video RAM consumption. We also run a task that calls a tool — not just a chat response — and verify the response time. Screenshots are kept as delivery documentation.
Your data stays on your hardware. Local deployment keeps records and configuration on your machine. The model can be swapped as better or lower-cost options emerge.
Pinned versions; no automatic updates. A version update is one of the most common causes of GPU fallback. We pin the version and validate any update before applying it — stability over chasing the latest release.
The value of local AI only materializes once it's confirmed to be running on the right hardware. That confirmation is what we deliver before handing over the keys.
FAQ · Frequently Asked Questions
- Q: How do I check whether my local AI model is using the GPU?
- For Ollama, run
ollama psand look at the PROCESSOR column. 100% GPU means fully on GPU; 100% CPU means the GPU is not being used; a split ratio means only part of the model fits in video RAM. - Q: I have a dedicated GPU — why is the AI still running on CPU?
- Common causes: outdated driver, an environment variable that has disabled GPU access, a model larger than available video RAM, or the GPU not detected by the software. Most of these produce no error message, so you have to check actively.
- Q: My local AI takes minutes per response. Is my hardware too weak?
- Not necessarily. In most cases we've investigated, the model was running entirely on CPU. Confirm GPU is being used before deciding to upgrade.
- Q: Is a bigger model always better?
- No. A model that doesn't fit in video RAM has part offloaded to CPU, which can make it slower than a smaller model that fits entirely on the GPU. Match model size to available video RAM.
- Q: How large should I set the context window?
- Ollama defaults to 4096 tokens. Increasing is fine, but each step up consumes more RAM; setting it too high can destabilize the whole machine. Increase gradually based on your hardware, not all at once.
This article is a Think BIG field notes publication, based on first-hand deployment experience. The views expressed do not represent any specific vendor or product. All named sources are linked inline; commands and default values are based on the official Ollama documentation current at time of publication.