THINK BIG
AI Agent Deployment
Thinking about abandoning your AI Agent? Ask these three questions first — and stop getting locked in
Most AI Agent abandonment has nothing to do with the tool itself. It comes down to three things nobody confirmed at the start: can you switch models, who owns the data, and can you predict the monthly cost? Get these three right upfront, and so-called "abandonment waves" will have nothing to do with you.
(Note: "Lobster" and "Hermes" are informal names used in the Taiwanese tech community for a category of AI Agent tools that can run continuously and operate the computer on your behalf. This article addresses deployment methodology for that type of tool, staying focused on the phenomenon without targeting any specific company or product.)
1. The Situation: What the conversation is really about
AI Agents — AI assistants that can take action and complete tasks on their own — have moved over the past year from early-adopter experiments into real workflows. Alongside that shift came a familiar pattern: install it, use it for two or three weeks, lose interest, leave it gathering digital dust. This "abandonment" isn't unique to Taiwan; the whole industry is grappling with it.
A few independently verifiable observations from published research: industry analysts project that by 2027, more than 40% of AI Agent projects may be abandoned or cancelled by businesses — and the cause is rarely the technology itself. The most common culprits are unclear governance and unexamined ROI assumptions.
One survey of senior executives found that roughly 80% of respondents were concerned about "over-reliance on a single AI vendor." Separate reporting notes that after months of accumulated configuration and context, switching AI Agent providers is far from a simple URL swap — and switching costs keep rising.
In short, abandonment usually comes down to two underlying problems: starting without a clear use case, then discovering the tool doesn't fit; and getting so embedded in a setup that you can't leave even when you want to. Whether the tool itself is good barely enters into it.
2. The Debate: Two legitimate perspectives
The optimistic view is that the core value of a continuously running AI Agent is precisely that you don't need to supervise it — it lives in your messaging app and on your computer, sends proactive reminders, reports on schedule, and learns your habits. For solo operators and small businesses, that's the equivalent of a tireless assistant who never clocks out.
The skeptical view centres on three concerns: lock-in risk (model, data, and configuration all tied to one vendor), cost unpredictability (pricing structures vary widely across the market — fixed fee, usage-based, or hybrid — and the difference matters over a year), and maintenance burden (getting it running and keeping it stable are two very different things).
These perspectives don't actually contradict each other. The optimists are talking about the value; the skeptics are talking about how to deploy without stepping on a landmine. Whether you end up abandoning the tool almost always comes down to whether you addressed those three skeptic concerns before you started.
3. What we learned from actually doing this
"Running" doesn't mean "running well." We deployed a local AI setup on one machine that looked fine — it answered questions — but every single response took one to four minutes, sometimes timing out entirely. After a long diagnostic detour, we found the model had been running entirely on the CPU. The GPU was doing nothing. Now we check GPU utilization before we change anything else.
One verifiable metric is all you need. For Ollama-based setups, ollama ps shows the PROCESSOR column (100% GPU, 100% CPU, or a split). The API also returns video RAM usage (size_vram). If video RAM usage is zero, it's running entirely on CPU — regardless of how normal things look. We screenshot this number before every delivery.
The reason is usually not "not enough hardware." Three root causes we've encountered, none of which require buying new equipment: the GPU wasn't detected (some integrated or hybrid graphics cards need extra environment configuration); the install path contained non-ASCII characters that crashed the detection process; the context window was set too large, which exhausted shared system memory and caused repeated crashes across the whole machine. AI memory is not "bigger is always better."
Nobody can promise perfect security. What we can do is check every layer: scan skills before installation, pin versions, audit again before going live, and keep keys in system config rather than in plain text. This isn't "we're safer than others" — it's "we've checked everything that can be checked."
4. Our recommendation: Three questions to answer before you deploy
Can you switch the underlying model? Is the tool tied to a single AI provider, or can you swap models freely? An architecture that lets you choose and switch means that when a better or cheaper model appears, you can move. That's the exit you want to have.
Who holds the data? Does it run on your own hardware, or on someone else's cloud? Local deployment keeps your data under your control. Cloud is also fine — but you should know exactly where your data sits before you commit.
Can you predict the monthly cost? Fixed fee, or does it scale with usage? A plan where you can estimate costs in advance and stay within budget serves long-term operations far better than one where the bill might surprise you.
One line to close with: abandonment is decided before deployment, not by the tool. Make sure you can switch models, that your data stays in your control, and that costs are predictable — and an AI Agent stops being a one-time experiment and starts being a genuine long-term collaborator.
What this architecture actually delivers
The three questions above are precisely what we lock down at the start of every deployment. The architecture we use delivers three concrete outcomes:
1. No single-model lock-in; API costs stay predictable. The underlying model is interchangeable — switch to a lower-cost model when saving budget, switch to a higher-capability model when the task demands it. The same assistant can cost dramatically different amounts from month to month, and you control it.
2. Your data stays on your hardware. Local deployment keeps records and configuration on your machine, not someone else's cloud.
3. Pinned versions, layered security review. No automatic updates; every skill goes through a security check before going live — stable, maintainable, and auditable.
This is why the so-called "abandonment wave" doesn't touch the people who deploy this way: model flexibility, data ownership, and cost predictability are built in from day one.
FAQ · Frequently Asked Questions
- Q: Is it normal to stop using an AI Agent after a few weeks?
- Very common. Industry projections suggest more than 40% of projects may be abandoned by 2027. But in most cases it's not the tool — it's that deployment started without clarifying the problem to solve, or whether the setup could be changed, or what the cost would look like.
- Q: How do I know if an AI Agent will lock me in?
- Check three things: can you choose and swap the underlying AI model; is the data on your machine or someone else's cloud; is the cost fixed or usage-based. Flexibility on all three keeps lock-in risk low.
- Q: Local or cloud — which suits me?
- Depends on the use case. If data residency matters, you want continuous operation, or you need predictable costs, local deployment makes more sense. For instant access without managing maintenance, cloud is more convenient. Both can be combined.
- Q: What should I watch out for with third-party skills or plugins?
- They typically have direct access to your files and accounts. Verify the source, run a basic security check, pin the version, and turn off auto-updates. Nobody can promise perfect security — the goal is to check every layer you can.
This article is a Think BIG field notes publication, based on first-hand deployment experience. The views expressed do not represent any specific vendor or product. All named sources are linked inline.