Computer vision — software that interprets images and video — has quietly become one of the most practical forms of AI a normal business can deploy. Not the headline-grabbing kind, but the sort that counts pallets, spots a defect, reads a label, or notices when someone steps into a danger zone. It works genuinely well on the right task, and it fails in specific, predictable ways on the wrong one. Knowing the difference is most of the value, and it is a difference the demos are designed to blur.
The tasks where it reliably works
Vision earns its place on jobs that are visual, repetitive, and precisely defined. Quality inspection is the classic case: a camera on a line checks each item for a defect faster and more consistently than a tired human at the end of a shift. Counting and measuring — pallets in a yard, items on a shelf, vehicles through a gate, the dimensions of a part — is another, because a machine does not lose count. Reading text and codes from labels, documents, and packaging turns a photo into structured data without manual typing. Safety monitoring watches for a missing helmet or a person in a restricted area. Inventory and stock checks read shelves or storage that a person would take hours to walk.
The common thread is a narrow, measurable question with a clear right answer: is this defective, how many are there, what does this label say, is someone where they should not be. Vision is strong when the question is that sharp and struggles the moment it is fuzzy. 'Is this weld cracked' is answerable. 'Does this look like a quality product' is a judgement wrapped in a hundred unstated criteria, and asking vision to make it is asking for a confident answer you cannot trust.
A worked example: catching a defect on a line
Picture a bottling line where a small share of caps seat crooked. A camera above the line photographs each cap, and a model trained on thousands of good and bad examples flags the crooked ones for rejection. In the pilot it looks magical — near-perfect, faster than any inspector. Then reality arrives in pieces. A new cap colour the model never saw. A reflection from a skylight at four in the afternoon that only appears in summer. A batch of bottles that sit a few millimetres higher in the new crates.
None of these is a flaw in the idea; each is a condition the system was never shown. The project succeeds or fails on whether someone anticipated them — photographed the new colours, shaded the camera, retrained when the crates changed — not on how clever the model was in week one. The lesson generalises: a vision system is a promise that the future will look enough like the images it learned from, and the engineering is in making that promise true as the world drifts.
The practical upshot is a rule of thumb: before trusting a vision system, ask what it has not seen yet. Every product variant, every season, every lighting condition, every new supplier of crates or caps is a gap until you have shown the system examples of it. A pilot that runs only through one clean month has not been tested; it has been flattered. The systems that survive contact with production are the ones whose owners kept feeding them the cases they got wrong, month after month, rather than the ones that scored highest on the first demo.
Where lighting and edge cases break it
Here is the honesty the demos skip. A vision system is only as good as the conditions it sees, and the real world is messier than the pilot. Lighting changes through the day and across seasons. Dust, glare, condensation, and a smudged lens degrade the image. The one product variant nobody thought to photograph shows up on the line. A part arrives at an angle the training data never contained. Each of these can turn a system that scored beautifully in testing into one that quietly misses defects in production.
This is not a reason to avoid vision. It is the reason to scope it to conditions you can control or at least anticipate, and to be ruthless about the edge cases before you trust it. A system that is right in good light and wrong in bad light is not 'mostly working' — it is a liability wherever the light is bad. Often the cheapest improvement to a vision project is not a better model but better lighting, a fixed camera position, and a shade over the lens. Controlling the scene is engineering too, and usually the highest-return kind.
On-device versus cloud
Where the analysis runs is a real decision with real consequences. Sending images to a cloud service is simplest to build and gives you the most computing power, but it needs reliable bandwidth, adds a round-trip of delay, and means your images leave your premises — which matters for both cost and privacy. Running the model on a device at the edge, next to the camera, keeps the images local, responds instantly, and works when the connection does not, at the price of more constrained hardware and a fiddlier deployment.
For a fast line that must react in milliseconds, or a site with weak connectivity, or footage you would rather never upload, edge is often the right call. For an occasional check where images can leave and latency does not matter, cloud is simpler. There is also a running-cost dimension people miss: analysing a continuous video stream in the cloud, frame after frame, day after day, can quietly become one of your larger bills, while the same work on a one-time piece of edge hardware is close to free after purchase. The point is to choose deliberately against the real constraints of the site, not to default to whichever a vendor sells.
The real work is data and labeling
The uncomfortable truth of almost every vision project is that the model is the easy part. The hard, slow, expensive part is collecting images that represent the real conditions and labeling them correctly — this one is a defect, this one is not, the object is here in this frame. A system learns from examples, and it can only recognise what it has been shown enough of, labeled consistently. A thousand photos all taken in perfect light teach it nothing about the bad light where it will actually fail.
Labeling is also where quiet inconsistency does the most damage. If two people labeling defects disagree about a borderline case — one calls a faint scratch a defect, the other does not — the model learns confusion and hands it back to you as unreliable output. Agreeing on the exact definition of the thing you are detecting, and labeling to that definition consistently, matters more than the volume of images. Plan for it honestly: gathering a representative set across the conditions you care about, and labeling it carefully, is the majority of the work and the foundation everything else stands on. A modest model on excellent, representative data beats an impressive model on thin, biased data every time.
Privacy and consent are not optional
Cameras that watch people carry obligations that cameras watching pallets do not. If your system sees employees, customers, or the public, you are processing personal data, and in the EU that comes with real requirements around lawful basis, transparency, and how long you keep footage. Facial features and identifiable people raise the stakes further. This is not a reason to abandon a safety or security use case, but it is a reason to design privacy in from the start — capturing only what the task needs, keeping it only as long as needed, and being clear with people about what is watching and why.
Good design often removes the problem rather than managing it. A system that counts people does not need to identify them; one that checks for a helmet does not need to store faces; much can be processed and discarded on the spot, so no identifiable footage is ever kept. Deciding this at the start is cheap. Retrofitting compliance onto a system that already hoovers up faces is expensive and sometimes impossible, and it is the kind of oversight that turns a useful safety tool into a regulatory liability.
Start with a pilot on one measurable task
The right way in is a single, narrow, measurable task, run as a pilot against real conditions before anyone commits to a rollout. Pick something with a clear right answer and a number you can check — this many defects caught, this count accurate to within one — and run the system beside your current method long enough to see how it behaves on the messy days, not just the clean ones. Measure it honestly, including the edge cases, and expand only once it has earned trust on the narrow task.
Be clear in advance about what 'good enough' means, because vision is rarely perfect and the honest question is whether it beats what you do now at a cost that makes sense. A system that catches most defects and flags the uncertain ones for a human can be transformative even though it is not flawless. If you are weighing whether vision fits a problem at all, start with an assessment of the task, the conditions, and the data you would need — that is where a project is won or lost, long before a model is trained.