All articles

How much does custom AI software cost?

The model is the cheap part. The real bill is data readiness, the un-glamorous software around it, and an inference cost ordinary software never had — here is how to think about it.

The honest answer to how much custom AI costs is that the model is almost never the expensive part. A demo that impresses a boardroom can be assembled in a week against a hosted model API for the price of a few coffees in tokens. The gap between that demo and something you can put in front of customers or trust with a business decision is where the real money lives — and most of it goes to work that looks nothing like artificial intelligence.

Data readiness is the biggest hidden cost

Before a model can do anything useful over your business, someone has to find, clean, and structure the data it will work from. In most companies that data is scattered across a CRM, a shared drive, five years of PDFs, and a database whose column names only one person still understands. It is duplicated, contradictory, and full of the exceptions every real business accumulates. Getting it into a state where a model produces reliable answers is often the single largest line item in an AI project, and it is the one buyers consistently underestimate because it is invisible in the demo.

This is not a step you can skip by throwing a bigger model at the problem. A capable model fed messy, ambiguous data produces confident, plausible, wrong answers — which is worse than no answer, because someone will act on it. The quality ceiling of an AI feature is set by your data long before it is set by the model.

There is a redeeming side worth naming. This work is not effort that evaporates once the model runs — a cleaned, well-structured, well-understood dataset is an asset in its own right that pays off in reporting, analytics, and every future feature you build. But it is genuine project cost, it comes first, and quietly assuming the data is ready is how a confident fixed-price quote turns into an uncomfortable conversation three months in.

Eighty percent of the build is ordinary software

The part everyone pictures — the clever model doing something surprising — is a small slice of the actual work. Around it sits the same software you would build for any serious product: authentication, a data pipeline that keeps the model's knowledge current, a user interface, logging, error handling, permissions, monitoring, and the plumbing that connects all of it to the systems you already run. None of that is glamorous, and all of it has to exist before the model's output is safe to rely on.

This is why an AI project is best budgeted as a software project that happens to include a model, not as a model that happens to need a little software around it. The intelligence is real, but it is a component. The product is everything you wrap around that component so a non-expert can use it every day without getting hurt.

It also explains why teams with strong software discipline tend to succeed with AI and teams without it struggle, regardless of how good the model is. If you cannot deploy, monitor, and roll back ordinary software reliably, adding a probabilistic component to the mix does not go well. The model amplifies whatever engineering culture it lands in — mature practice makes it dependable, and its absence makes it a liability that is now harder to debug.

Evaluation and guardrails are not optional extras

Ordinary software is deterministic — the same input gives the same output, so you can test it and move on. A model is probabilistic. It can be right ninety-five times and confidently wrong the ninety-sixth, and you often cannot tell which from the answer alone. That changes what testing means. You need an evaluation harness: a growing set of real examples with known-good answers that you run the system against every time you change anything, so you can measure whether a change made it better or quietly worse.

Alongside that sit guardrails — the checks that stop the system doing something harmful, leaking data it should not, or answering a question it has no basis to answer. Building and maintaining these is real engineering effort, and it is the effort that separates a toy from something a regulated business can actually deploy. Skipping it does not save money; it defers the cost to the day the system embarrasses you in front of a customer.

Model API versus running your own

One real cost decision is whether to call a hosted model through an API or run an open model on your own infrastructure. Using an API is far cheaper and faster to start: no hardware, no operations team, and you pay per use. For most companies most of the time, it is the right first choice. Self-hosting becomes worth considering when data cannot leave your environment for legal or contractual reasons, when your volume is high enough that per-call pricing overtakes the cost of running your own hardware, or when you need a level of control an external provider will not give you.

Self-hosting trades a predictable per-call fee for a fixed, and not small, cost in GPUs and the people who keep them running. It is a genuine option, not a default. The right answer depends on your volume, your data sensitivity, and whether you have anyone who wants to operate inference infrastructure — and that is a decision worth making deliberately, not by habit.

It is also worth knowing this is a spectrum, not a switch. Between the public API and running your own hardware sit managed options — providers that host an open model for you, or run a dedicated instance in a region you choose. For a company with real data-residency concerns but no wish to operate GPUs, that middle ground is often the pragmatic answer, and the sensible choice usually lands further toward the managed end than engineers instinctively reach for.

Prototyping is cheap; reliability is not

Here is the pattern that surprises buyers most. Getting to a working prototype — something that does the impressive thing most of the time — is genuinely cheap and fast now. Getting from there to something reliable enough to trust with real work is the long, expensive stretch. The last ten percent of reliability can cost more than the first ninety, because it means handling every edge case, every malformed input, every way a user will use the thing you did not anticipate, and every failure mode you would rather not think about.

This is not a reason to avoid AI. It is a reason to be honest in the budget about which stage you are paying for. A prototype proves the idea is possible. Production proves it is safe. Confusing the price of the first for the price of the second is how AI projects run over.

A practical consequence follows for how you buy. Treat the impressive prototype as evidence that the idea is worth pursuing, not as ninety percent of a finished product with only polish remaining. The polish is the product. Planning the budget as though the demo were nearly done is the most common way an AI initiative arrives late and over cost, with everyone surprised that the last stretch took longer than the first.

The running cost ordinary software does not have

Traditional software costs money to build and comparatively little to run — you pay for some servers and they largely idle. AI is different, and this catches people out. Every time the model answers, it costs money: tokens if you use an API, GPU time if you self-host. A feature that gets popular gets more expensive to run, not less. That ongoing inference cost is a permanent line in your operating budget, and it scales with usage rather than sitting flat.

This is not a flaw; it is just the shape of the thing. But it means you cannot evaluate an AI feature on build cost alone. A feature that is cheap to build and expensive per use can quietly become your largest cloud line once it succeeds — so the running cost belongs in the business case from the start, not as a surprise on the first full bill.

The good news is that this cost is manageable once you see it clearly. Much of it can be tuned — caching repeated answers, routing simple requests to a smaller cheaper model and reserving the expensive one for hard cases, and setting sensible limits so a single runaway process cannot generate a shocking bill. None of that is exotic, but it only happens if someone owns the running cost as a first-class concern rather than discovering it after launch.

Start narrow, buy certainty

The way to keep an AI budget honest is to refuse to price the whole thing up front, because at the start neither of us knows enough to price it well. Instead, start with one narrow, valuable use case — a single process, a single kind of question — and build it end to end, through the data work, the guardrails, and the evaluation, to something real people use. That first slice tells you what the data actually costs to prepare, how reliable the model can get on your problem, and what a query really costs to serve. Those three numbers turn every later estimate from a guess into arithmetic. The narrow pilot is not a way to save money on the ambition — it is how you buy the certainty that makes the ambition fundable.

Seen this way, the pilot is a cost-control tool as much as a technical one. You spend a bounded, known amount to remove the biggest unknowns before anyone commits to the full ambition. If it shows the data is cleaner than feared and the model handles your problem well, you scale with real confidence. If it shows the opposite, you learned that for a fraction of what learning it the hard way would have cost. Either outcome is worth the price of the pilot.

Wondering what an AI feature would really cost you?

A short, fixed-fee assessment scopes one narrow use case, checks your data, and turns a vague ambition into a costed first phase with a real running-cost estimate — before you commit a budget.

Get a scoped AI assessment