Every product team is under some version of the same pressure right now: add AI, or look like you are falling behind. It is a real pressure and it produces a predictable, wasteful pattern — a chatbot bolted onto a product that did not need one, a feature shipped because the technology is impressive rather than because it does a job. By 2026 the novelty has worn off enough that we can be plain about it: an LLM is a normal engineering component with specific, well-understood failure modes. Used where those modes are tolerable, it is genuinely transformative. Used where they are not, it makes your product worse in ways that are hard to see until a customer is harmed.
Start from the job, not the technology
The first mistake is deciding to add AI and then hunting for somewhere to put it. That is backwards, and it reliably produces features nobody asked for. The right starting question is the one you would ask about any feature: what job is the user trying to get done, and where are they currently struggling? Only then do you ask whether AI is the best tool for that specific job — often it is, sometimes a boring database query or a well-designed form beats it outright.
This matters because AI is not free to add. It brings cost, latency, and a whole new class of failure. A feature has to earn all of that by doing a job meaningfully better than the alternatives, not merely by existing. If you cannot name the job in a sentence, you are not ready to build the feature — you are shopping for a use case, and users can tell.
A useful discipline here is to imagine the feature without the AI. If a rules engine, a search index, or a short form would do the job acceptably, that is very often the answer — cheaper to run, faster, and predictable in a way a model is not. AI earns its place when the job is genuinely beyond what those tools can do: when the input is unstructured language, when the space of valid answers is too large to enumerate, when a human would use judgement. If a simpler tool clears the bar, use it and spend the AI budget where only AI will do.
Where AI genuinely helps
There is a clear shape to the tasks where language models shine, and it is worth internalizing because it predicts success far better than intuition does. AI is strong at tasks that are fuzzy, forgiving, and language-shaped. Summarizing a long document. Extracting structured fields from messy text. Drafting a first version of something a human will edit. Classifying free-form input into categories. Semantic search, where the user means something the exact keywords do not say. Assisting a human who stays in control of the outcome.
What these share is that the cost of a mistake is low and a human is positioned to catch it. A draft that is slightly off is still a useful starting point. A summary that misses a nuance is still faster than reading the whole thing. In these jobs the model does not have to be perfect to be valuable — it has to be helpful more often than not, with a person able to correct the rest. That is a bar a good LLM feature clears comfortably.
It is also worth noticing that the strongest AI features tend to be assistive rather than autonomous. The model drafts and the person sends; the model suggests a category and the person confirms; the model surfaces candidates and the person chooses. Framed that way, the feature borrows the model's speed and breadth while keeping the human's judgement in charge of the outcome — and users trust it more precisely because they remain in control. Autonomy is where the stakes and the failure modes both climb fastest.
Where it hurts
The mirror image is just as clear, and it is where most AI features go wrong. AI is a poor fit for tasks that are precision-critical or deterministic — where there is one correct answer, the answer must be exactly right, and a wrong one causes real harm. Calculating a customer's invoice. Deciding whether a transaction is fraudulent with no human review. Anything where a confident, plausible, wrong answer is worse than no answer at all.
That last point is the crux. A traditional system that cannot answer returns an error, and an error is honest — it tells you it failed. A language model rarely does that. It produces a fluent, confident, wrong answer that looks exactly like a right one, and confidence is precisely what makes it dangerous in a precision-critical setting. If your users will trust the output and act on it, and being wrong costs them money, safety, or trust, that is the place not to hand the decision to a model. Use it to assist the human making the call, not to make the call.
Design for the reliability problem
Language models make things up. They call it hallucination, and it is not a bug that a better model quietly retires — it is an inherent property of how these systems work, and by 2026 the honest position is that you design around it rather than waiting for it to disappear. That means building the feature to expect wrong output, not to be surprised by it. Ground the model in real data so its answers are drawn from your sources rather than invented. Give it a way to say it does not know instead of forcing a guess. Show the user where an answer came from so they can check it. And keep a human in the loop wherever the stakes justify it — the model proposes, the person disposes.
The point is not to make the model perfect, which is not on offer. The point is to build a system that stays useful and safe even when the model is wrong — because it will be, and a design that assumes otherwise fails in production in front of a customer.
Cost, latency, data — the constraints that are easy to forget
An AI feature has running costs that scale with use, in a way most software features do not — every call to a capable model costs real money, and a feature that is cheap in a demo can be alarming at scale. It is also slow: a model that takes several seconds to respond changes what interactions are even sensible, and a user waiting on a spinner for a task that used to be instant will not thank you for the AI behind it. Design for both from the start rather than discovering them after launch.
Then there is data. Sending customer information to a model, especially a third-party one, raises exactly the privacy and residency questions any data processing does — more so under GDPR in the EU. You need to know what leaves your system, where it goes, and whether you are permitted to send it. This is not a reason to avoid AI; it is a set of engineering and legal decisions to make deliberately, before the feature ships, rather than a surprise you discover in an audit.
There is a design lever that addresses cost and latency together, and it is chronically underused: not every task needs the most capable, most expensive model. Much of the value comes from routing the easy cases to a smaller, faster, cheaper model and reserving the expensive one for the hard ones, or from caching results that recur. Treating one giant model call as the only tool is how a feature that works becomes a feature you cannot afford to leave on.
Build on evaluations, not vibes
Traditional software is tested against expected outputs — you know what correct looks like and you assert it. AI features resist that because the same input can produce different valid outputs, and correctness is often a judgement rather than a match. The failure mode this creates is building on vibes: a few impressive demos, a good feeling, and a feature shipped with no real measure of how often it is actually right. That feature will degrade, or a model update will shift its behaviour, and you will have no way to know until users complain.
Evaluations also change how you ship. With a real quality number in hand, a prompt tweak or a model swap stops being a leap of faith and becomes a measured change — you run it against the same cases and see whether the number moves. That is what lets an AI feature improve steadily instead of drifting, and it is what turns a provider's model update from a source of dread into something you can absorb on purpose. Without it you are flying blind, redecorating prompts and hoping.
The discipline that replaces vibes is evaluations — a repeatable set of real cases with a defined notion of a good answer, run against the feature so you have an actual number for its quality and can watch that number as you change prompts, models, or data. It is unglamorous, and it is the single practice that separates AI features that hold up from ones that quietly rot. Treat an LLM feature as what it now is — normal engineering with specific failure modes — and the whole thing becomes tractable. If you are weighing an AI feature, the most useful first step is a short conversation to sort the jobs where it will genuinely help from the ones where it will quietly hurt.