Most software is bought for what it does. It is kept for whether it stays up. A feature that dazzles in a demo is worthless the morning it will not load, and the phone call you get is never about the clever part — it is about the login page returning an error while your customers wait. Reliability is the quiet feature nobody asks for until it is missing, and it is worth understanding before you need it rather than during the outage.
What uptime and the "nines" really mean
Uptime is the share of time your software is available and working. It is usually quoted in nines. 99.9% — "three nines" — sounds almost perfect, but it allows roughly forty-three minutes of downtime a month. 99.99% — "four nines" — allows about four minutes a month. The gap between those two numbers looks tiny on paper and is enormous in practice, because the last minutes of downtime are the hardest and most expensive to eliminate.
The number that matters is not the one on a marketing page but the one measured against your reality. A batch system that runs overnight can be down for an hour at noon and no one notices; a checkout that fails for four minutes on your busiest day is a serious incident. Before you fixate on nines, ask a plainer question: which minutes of the day actually matter for us, and what does it cost us when the software is down during them? That framing turns an abstract percentage into a business decision.
It also helps to know that downtime is rarely one long outage. It is usually many short ones — a few seconds here, a minute there — that add up to the monthly total, plus the occasional bad hour that dominates it. Two systems can share the same headline uptime and feel completely different: one that blinks for ten seconds twice a day is annoying, while one that is solidly up for a month and then down for forty minutes in a single stretch during business hours is a crisis. When you set a target, set it against the shape of failure your business can least afford, not just the total.
You cannot fix what you cannot see
The first requirement of reliability is not more servers. It is knowing. Monitoring is the software watching itself — response times, error rates, whether the important paths still work — and alerting is what turns a worrying pattern into someone's phone ringing before customers notice. The failure mode you want to design out is the one where your first alert is a customer email, because by then the outage has been running for however long it took someone to complain.
Good monitoring watches the things your users care about, not just whether a machine is powered on. A server can be perfectly healthy while the checkout it hosts is silently failing, so the useful check is "can a customer actually complete a purchase right now," posed continuously and answered without a human in the loop. When you evaluate a partner, ask how they would find out you were down. If the honest answer is that they would wait for you to tell them, you have learned the true state of their monitoring.
No single point of failure
Every reliable system is built on the same principle: nothing important depends on a single thing that can fail. If one server going down takes the whole application with it, that server is a single point of failure, and its failure is not a possibility but a certainty on a long enough timeline. Redundancy means running more than one of each critical part, so the loss of any one is absorbed rather than fatal — two servers where the second carries the load if the first dies, a database that keeps a live copy ready to take over.
Redundancy costs money and adds complexity, which is exactly why it is worth doing deliberately rather than everywhere. You do not need every component duplicated; you need the ones whose failure would stop the business duplicated, and the rest designed to fail without taking the important parts down with them. The engineering skill is knowing which is which. A partner who cannot point to your single points of failure has not looked for them, and the ones you have not looked for are the ones that take you down.
The subtle single points of failure are the ones outside the obvious server list. A single database everything writes to, one payment provider with no fallback, a certificate that quietly expires, one person who is the only one who knows how a deployment works — each is a dependency whose failure stops you, and none of them is a spare server away from safe. Mapping them honestly is uncomfortable, because the list usually includes something nobody wants to own. That discomfort is the point: the failures that hurt most are the ones the organisation had quietly agreed not to look at.
A backup you have never restored is a hope, not a plan
Backups are the part everyone agrees on and few people test. Taking a backup is easy; the question that decides whether it was worth taking is whether you can actually restore from it, how long that takes, and how much data you lose in the gap. Disaster recovery is the plan for the bad day — a data centre lost, a database corrupted, a mistake that deletes what it should not — and a plan that has never been rehearsed is not a plan. It is an assumption, and the middle of a crisis is the worst possible time to discover it was wrong.
Two plain numbers make this concrete. How much data can you afford to lose, measured in time — a minute, an hour, a day? And how long can you afford to be down while you recover? Those two answers drive the entire design, and they are business decisions, not technical ones. A backup taken hourly and never restored gives false comfort; a backup taken hourly and restored in a routine drill gives you a number you can trust. Insist on the drill.
Someone has to answer the phone
Software does not fail politely at eleven in the morning. It fails at the worst time, and reliability depends on what happens in the first fifteen minutes after it does. On-call means a named person is responsible for responding when an alert fires, at any hour, with the access and authority to act. Incident response is the practised routine around that — who is called, how the problem is contained before it is fully understood, how customers are kept informed, and how the team makes sure the same failure does not happen twice.
The tell of a mature team is not that they never have incidents; it is that their incidents get shorter and rarer, because each one ends with a calm look at what happened and a change that removes the cause. A team without on-call is telling you that when something breaks outside office hours, it stays broken until office hours. For some software that is a perfectly reasonable choice. For a system your customers use at nine on a Sunday evening, it is a decision you should make on purpose, not discover by accident.
The cost of the last nine
Reliability does not scale linearly with effort, and this is the fact that reframes the whole conversation. Getting from unreliable to 99.9% is mostly good basic practice — monitoring, sensible redundancy, tested backups. Getting from 99.9% to 99.99% is a different kind of project: it demands eliminating single points of failure you were previously willing to tolerate, automating recovery so no human delay is in the path, and testing failure deliberately. The step looks small — one more nine — and the effort behind it is often several times larger than everything that came before.
This is why the right question is never "how do we get maximum uptime." It is "how much uptime is worth paying for," and the honest answer is different for a hospital system than for an internal reporting tool. Chasing nines you do not need spends money that would do more good elsewhere; skimping on nines you do need saves money you will hand back many times over in a single bad outage. The engineering is in matching the two.
Deciding how much reliability you actually need
An SLA — a service level agreement — is where a reliability target stops being an aspiration and becomes a commitment with consequences. It states the uptime you can expect and what happens if it is not met. A meaningful SLA is backed by the architecture and the on-call rota that make the number achievable; an SLA written to win a deal, with nothing behind it, is a number someone will quote back to you angrily during your first serious outage. Ask what the promise is built on, not just what it says.
So decide deliberately. Map the parts of your software to what an outage of each would actually cost — in lost revenue, in trust, in the hours your team spends firefighting — and buy reliability in proportion. Some systems justify four nines and the expense that comes with them; many are served perfectly well by three, with a clear recovery plan for the rare bad day. The goal is not the highest number. It is the right number, chosen with eyes open, so that you are neither paying for reliability you will never use nor discovering in production the reliability you forgot to buy.
A useful exercise is to write down, for each part of your software, the sentence you would have to say to a customer during an outage of it — and how long you could say it before the damage was lasting. Some sentences are survivable for an afternoon; some are not survivable for five minutes. That sentence, not a percentage, is what tells you where to spend. Reliability bought this way lands where it earns its keep, and the parts that can quietly wait are allowed to quietly wait.