You have decided to do something with AI. That is a direction, not a decision. The technology is now cheap to try and easy to demo, which is exactly why so many companies end up with an impressive prototype that never earns back the time it cost. The projects that pay off are not the ones with the cleverest model. They are the ones that started from the right problem. Choosing that problem well is most of the work, and it happens before a single line of code.
Start from a problem, not from the technology
The wrong way to start is with the tool in hand, looking for somewhere to point it. That produces demos in search of a purpose — a chatbot nobody asked for, a summariser that saves nobody any time. The right way starts with a problem that already hurts, phrased without the word AI in it at all. Where do your people spend hours on work that is repetitive and mechanical? Where does a queue build up because a human has to read, sort, or classify something before anything else can happen?
The best first candidates share a shape. They are high-volume, so automating them frees real hours rather than a rounding error. They are repetitive, so the task looks similar each time. And they are tolerant of the occasional wrong answer, because no model is perfect and your first one certainly will not be. A problem with all three properties is a gift. A problem with none of them will punish you no matter how good the technology gets.
One useful habit: describe the problem to someone outside the company in a single sentence, and watch whether AI appears in it. It should not. If you cannot state the problem without naming the solution, you are still reasoning from the technology, and you will end up building something interesting rather than useful. The problems worth solving are boring to describe and expensive to live with — a queue that never clears, a report that eats a day every week, a decision that always waits on a person who is always busy.
The value test: does it save real time or unlock revenue
Once you have a candidate, ask what it is actually worth. Not in the abstract — in hours, in euros, in a bottleneck removed. If a task takes your team forty hours a week and AI can take a serious bite out of that, the value is legible to everyone, including the person approving the budget. If the benefit is vague — faster, smarter, more modern — you do not have a value case, you have a mood.
Value comes in two honest forms. The first is cost taken out: work that no longer needs a person, or a person freed to do something only a person can do. The second is revenue unlocked: something you could not offer before, or could not offer at the speed the market now expects. Both are real. What is not real is value that lives entirely in a slide. If you cannot point to the hour saved or the sale enabled, keep looking.
Be wary, too, of value that is real but not yours to capture. A use case can save an enormous amount of time in theory and free no one in practice, because the hours it removes are scattered across many people in small slivers rather than concentrated where they can be redeployed. Ten minutes saved for fifty people is fifty people with ten spare minutes, not a role you can reassign. The value that shows up on a budget is the value that lands in one place large enough to act on.
The feasibility test: is the data there, is a wrong answer survivable
A valuable use case you cannot build is not a use case. Feasibility comes down to two questions, and they are the ones most often skipped. First: does the data exist, in a form you can actually reach? An AI system learns from and acts on data, and if the examples it needs are locked in someone's inbox, scattered across formats, or simply never recorded, the project stalls before it starts — not on the model, on the plumbing.
Second: what happens when the system is wrong? Because it will be. If a wrong answer is caught by a human before it does harm, or costs a small correction, you can ship and improve. If a wrong answer goes straight to a customer, a regulator, or an irreversible action, the bar is far higher and the first project is far riskier. The safest early wins put a person between the model and the consequence. The model drafts; a human approves. That single arrangement makes a surprising number of use cases feasible that would otherwise be reckless.
Why customer-facing and precision-critical is a bad first bet
The instinct is to aim AI at the most visible problem — the thing customers see, the number the board watches. Resist it as a starting point. A customer-facing, precision-critical use case is exactly where the tolerance for a wrong answer is lowest and the cost of a public mistake is highest. It is the hardest possible place to learn how AI behaves in your business, and you will be learning in front of the people you least want to disappoint.
Start where the stakes are lower and the feedback is faster: internal work, back-office tasks, a draft a colleague checks before it goes anywhere. You build the same muscles — data, evaluation, the honest sense of where the model helps and where it does not — without betting the brand on your first attempt. Once you have earned that judgment internally, the customer-facing use case becomes a considered second move rather than a gamble.
A convincing demo is not a proven use case
AI demos beautifully. That is precisely the danger. A model that dazzles on a curated example in a meeting is telling you almost nothing about how it behaves on the messy, adversarial, edge-case-ridden reality of your actual work. The gap between a demo and a dependable system is where most of the disappointment in AI lives, and it is invisible at exactly the moment a budget gets approved on the strength of a slick five-minute showing.
When you evaluate a candidate, discount the demo and interrogate the hard cases. What does it do with the input that is malformed, ambiguous, or unlike anything in the examples? How often is it confidently wrong, which is far more dangerous than being uncertain? A use case that looks easy in the demo and collapses on the long tail of real inputs is worse than one that looks modest and holds up, because the first sets an expectation the second never made.
This is not cynicism about the technology; it is respect for it. The teams that get durable value from AI treat every impressive demo as a hypothesis to be tested, not a result to be celebrated. Choose the use case that survives the hard questions, not the one that gives the best demo.
Score your candidates before you pick one
You will usually have several ideas, not one. Rather than argue them by conviction, score them on the same axes. For each candidate, rate the value — hours or euros, concretely — and rate the feasibility — data availability and the cost of being wrong. A candidate that is high value and high feasibility is your first project. High value but low feasibility is a later project, once the data is in order. Low value but easy is a distraction that will consume attention it does not deserve. Low on both is a no.
The point of scoring is not the number; it is the conversation it forces. When two stakeholders disagree, they are almost always disagreeing about value or about feasibility without naming which. Putting both on the table separates the argument you can settle with facts from the one you can settle with a small experiment.
Resist the urge to score in private and announce the winner. The value of the exercise is that it is shared: when finance, operations, and engineering rate the same candidates on the same two axes, the disagreements surface early and cheaply, on a page, instead of late and expensively, in a half-built project. A use case everyone rates high is one that will have sponsors when it needs them — and that matters as much to whether it ships as any technical property does.
Run a small pilot before you commit
No amount of scoring replaces evidence. Before you commit a real budget, run a narrow pilot on the top candidate — a few weeks, a fixed scope, a clear question: does this work well enough, on our real data, to be worth building properly? A pilot is not a proof of concept that lives forever in a demo. It is a decision tool with a deadline. It either earns the next phase or it saves you from a project that looked good only on paper.
Keep the pilot small on purpose. Pick one workflow, one team, one measurable before-and-after. Instrument it so you can tell whether the model is genuinely better than what you do today, not merely novel. And decide in advance what result would make you stop — the discipline that separates a real pilot from an expensive way of talking yourself into something.
Watch out for the pilot that quietly becomes permanent. A prototype wired up to impress in a demo has a way of surviving into production, because tearing it down feels like waste — and then you are operating something that was never built to be operated. Decide before you start whether a successful pilot will be rebuilt properly or promoted as it stands, and be honest that a pilot built to answer a question is rarely the same thing as a system built to run for years. Choose the problem well, prove it small, and the rest of the AI project stops being a leap of faith and becomes ordinary, sequenced engineering.