You ask two people for last month's revenue and get two different numbers. Finance pulls one figure from the accounting system, sales pulls another from the CRM, and the argument about which is right takes longer than the report itself. Meanwhile the monthly board pack takes three days to assemble by hand, and half of that is copying between spreadsheets. This is not a dashboard problem. It is that your company has no single place where data is trusted, and building that place is what a data warehouse is for.
The real problem: numbers that never match
Scattered data is the default state of any company that has grown past a handful of tools. The CRM knows about deals, the billing system knows about invoices, the support desk knows about tickets, and the web analytics knows about traffic. Each of those systems is a source of truth for its own job and only its own job. The moment you want a number that spans two of them — revenue per customer, cost to serve, churn by acquisition channel — there is no system whose job that is.
So someone exports, someone pastes, someone builds a spreadsheet that becomes load-bearing, and that spreadsheet quietly disagrees with the next one because they were pulled on different days with different filters. The numbers do not match because nothing was ever responsible for making them match. A warehouse is the thing you make responsible.
The cost of that gap is easy to underestimate because it does not show up as a line item. It shows up as senior people spending days a month assembling reports instead of acting on them, as decisions delayed until the numbers can be reconciled, and — most corrosively — as a slow erosion of trust in every figure the company produces. Once people have been burned by a report that turned out wrong, they start keeping their own private version of the truth, and now you have the original problem multiplied by everyone who stopped believing the shared one.
What a warehouse actually is
Strip away the vendor language and a data warehouse is one database that every other system feeds into. Data flows in from each source, gets cleaned and reshaped into a consistent structure, and then lives in one place that your reports and dashboards read from. Nothing writes to it except the pipelines that load it. It is deliberately not the system anyone runs their day-to-day work in — it is the system you ask questions of.
That separation is the whole point. Your CRM is built to be fast at showing one salesperson their deals, not at summing four years of orders across every region. A warehouse is built for the opposite: it is slow to change and fast to ask sweeping questions of. When people say a report took three days, what they usually mean is that they were doing a warehouse's job by hand, every month, without one.
A warehouse also keeps history in a way source systems rarely do. Operational systems care about now — the current balance, the open tickets, this month's pipeline — and they overwrite the past as it stops being current. A warehouse remembers, so you can ask how a number looked last quarter and get the same answer you would have got then. That memory is quietly one of the most valuable things it gives you, because most real business questions are about change over time, not a single snapshot.
Pipelines, ELT, and where the work lives
Getting data from a source into the warehouse is a pipeline. The old approach, ETL, extracted the data, transformed it into its final shape on the way, and then loaded it. The now-common approach, ELT, loads the raw data first and transforms it once it is inside the warehouse. The reordering matters more than it sounds: with ELT the warehouse keeps a copy of the raw data, so when a definition changes — and it will — you reshape from the raw copy instead of re-extracting from a source that may no longer have last year's data.
You do not need to memorise the acronyms. What you need to hold onto is that the hard, valuable work is the transformation: turning ten systems' idea of a customer into one, deciding how a refund reduces revenue, deciding what a currency conversion uses as its rate. That logic is your business encoded as data, and it is worth writing down carefully once rather than re-deriving it in every spreadsheet forever.
The other half of a pipeline is the unglamorous operational part: how often it runs, what happens when a source is briefly unavailable, and how it loads only what changed instead of re-reading everything every night. Freshness is a real decision, not a default — some numbers genuinely need to be current within minutes, most are fine the next morning, and paying for minute-by-minute freshness on data nobody looks at before nine is a common and quiet waste. Decide what each report actually needs and build to that, not to the fastest option available.
Agree on definitions before you agree on tools
The most expensive arguments in analytics are not about technology. They are about what a word means. Does revenue include tax? Does it count the day the deal is signed or the day the invoice is paid? Is a customer who cancelled and resubscribed one customer or two? An active user — active in the last day, week, or month? Until these are settled and written down, every dashboard is just a well-formatted opinion.
A warehouse forces the question, and that is a feature. Because everything reads from one modelled layer, a definition can only exist once. When you decide that revenue is recognised on invoice and net of refunds, that decision lives in one transformation and flows into every report at the same time. The discipline is agreeing; the warehouse just makes disagreement visible instead of letting it hide in a hundred separate exports.
This is why a warehouse project is never only a technical project. It surfaces questions the business has been quietly answering inconsistently for years, and it forces someone to make a call. That is uncomfortable, and it is also the single most valuable side effect of the work. A company that has agreed, in writing, on what its own core numbers mean is in a stronger position than one with a faster dashboard and no such agreement.
Build it yourself or buy the managed pieces
You are not choosing between a warehouse and no warehouse — you are choosing how much of it to run. The storage-and-query engine is almost always something you rent rather than build; managed cloud warehouses have made running your own database engine hard to justify for this. The pipelines that load common sources are increasingly something you buy too, because a connector to a popular CRM is a solved problem and rebuilding it is a poor use of your engineers. What is genuinely yours, and worth building, is the transformation layer — the business logic that turns raw data into your company's actual definitions.
The honest tradeoff is between speed and fit. Managed tools get you a working pipeline in days and are the right default for standard sources. Custom work earns its cost only where your data or your questions are unusual enough that no off-the-shelf connector understands them. A good build spends its custom effort exactly there and buys everything else, rather than treating a warehouse as a chance to hand-craft parts that a vendor already solved better and cheaper than you will.
Start with the questions, not the whole warehouse
The failure mode is trying to model everything before anyone gets an answer. A warehouse that ingests all forty of your systems and produces nothing useful for six months is a project that gets cancelled in month five. Turn it around: name the three questions the business most needs answered — the ones people currently spend days assembling by hand — and build only the slice of the warehouse that answers them.
That slice touches maybe three sources instead of forty, needs a handful of definitions instead of hundreds, and produces a report someone actually uses within weeks. It also proves the approach to the people paying for it, which is how you earn the room to add the next slice. A warehouse is grown, not delivered; the first harvest should come early.
Starting from the questions also protects you from the most seductive mistake in this whole field — building a beautiful, complete model of your data that answers questions nobody asked. Completeness is not the goal. A report that changes a decision is the goal, and you get there fastest by working backwards from the decision to the smallest set of data that informs it.
Who keeps it running afterwards
A warehouse is not a project that ends. Sources change their formats, a new system gets adopted, a definition shifts when the business changes how it sells. If no one owns the pipelines, they break quietly and people drift back to their private spreadsheets — and you have paid for a warehouse to end up exactly where you started. Decide before you build who watches the loads, who arbitrates when a definition needs to change, and who says no when someone wants to bolt reporting logic onto a source system instead of the warehouse.
That ownership is cheap compared to the alternative, which is a trusted system slowly becoming untrusted. It does not take a large team — often it is one person with clear responsibility and the authority to make definitional calls stick. What it cannot be is nobody. A warehouse that no one owns has the same lifespan as the enthusiasm of whoever built it, and enthusiasm is not a maintenance plan.