All articles

GDPR and custom software: building data protection in from day one

GDPR is easiest and cheapest when it shapes the architecture rather than being bolted on. This is engineering guidance, not legal advice.

Most teams meet GDPR as a document — a policy someone in legal wrote, a checklist to sign before launch. By then the architecture is set, the data model is fixed, and compliance becomes a scramble of workarounds bolted onto a system that was never shaped for it. It does not have to go that way. Handled well, data protection is an architecture decision made early, and early it is cheap. One note before we start: this is engineering guidance from people who build systems, not legal advice. For how the rules apply to your specific situation, talk to a qualified data-protection advisor — this article is about how to build so that whatever the law asks, the system can answer.

By design and by default is an architecture decision

The phrase at the heart of GDPR is data protection by design and by default, and the temptation is to read it as a compliance slogan. It is not. It is an instruction about how to build. By design means protection is considered when you draw the system, not after it runs. By default means the privacy-protective setting is the one that ships — data is not collected, shared, or retained unless there is a reason and someone chose it.

In practice this changes early decisions. Where does personal data enter the system, and does it need to. Which services touch it, and can that be fewer. How is it separated from data that is not sensitive. These are architecture questions, and they are nearly free to answer at the whiteboard and expensive to answer once the data is spread across a dozen tables and three integrations. The whole argument of this article is that timing, not effort, is what makes compliance cheap or ruinous — a decision made in a design meeting costs an hour, and the same decision forced onto a live system can cost a quarter.

Minimize, name a purpose, set a retention

Three ideas sit at the core of the regulation and they are genuinely useful engineering constraints, not bureaucracy. The first is minimization: collect only what you actually need for the task at hand. The instinct to capture everything because it might be useful someday is exactly what the rule pushes against — and it is also good engineering, because data you never hold cannot leak, cannot be misused, and never has to be protected, migrated, or explained.

The second is purpose. Every piece of personal data should exist for a stated reason, and using it for something else later is not a small thing. If you collect an email to send a receipt, quietly repurposing it for marketing is the kind of drift the regulation exists to stop. The third is retention: data should not live forever by default. Most systems keep everything because deleting is work no one scheduled, and old records quietly accumulate into a liability no one is watching. Deciding up front how long each kind of data lives, and building the mechanism to enforce it, turns a growing risk into a bounded one. All three are cheaper as design rules than as later corrections.

Access control and audit

Once you hold personal data, two questions follow: who can reach it, and can you prove who did. Access control means not everyone in the company sees everything. An engineer debugging a payment issue rarely needs the customer's full history; a support agent needs what the ticket requires and no more. Least privilege — each person and each service granted only what its job demands — is both a security principle and a data-protection one, and it is far easier to build in than to retrofit onto a system where everyone already has broad access and no one remembers why.

Audit is the other half. When personal data is viewed, changed, exported, or deleted, that should leave a trail. Not to police your own people, but because being able to answer who accessed what, and when, is part of taking the data seriously — and it is the difference between knowing what happened after an incident and guessing under pressure. Logging done from the start is quiet infrastructure that costs almost nothing; added afterward it means threading instrumentation through code that never expected it, touching far more of the system than anyone estimated.

Where the data lives, and who else touches it

For an EU company, where personal data physically resides is a real design input. Keeping EU residents' data within the EU is the simplest posture — data hosted in an EU region, on infrastructure whose location you can state plainly to a customer or a regulator. It is not the only lawful arrangement, but transfers outside the EU carry conditions, and the cleanest way to avoid a tangle is often to not create one in the first place.

Then there is everyone else who touches the data on your behalf. The cloud host, the email service, the analytics tool, the payment provider — under GDPR these are processors, and the ones they in turn rely on are sub-processors. You remain responsible for the personal data even when it flows through them. That means knowing who they are, where they operate, and having the right agreements in place. Every third-party service you wire in is a data-protection decision, not just a technical one, and the time to make it is when you choose the service — not when someone asks for the list and you discover no one kept one.

Erasure and export are features you build

GDPR gives people rights over their data, and two of them land squarely on engineering. The right to erasure means a person can ask you to delete their personal data, and you have to be able to actually do it. That sounds simple until you look at a real system where a user's data is scattered across the main database, backups, logs, a search index, a data warehouse, and three external services. If deletion was not designed for, honoring the request becomes an archaeology project that no one is sure came out complete.

The right to access and portability is the mirror image: a person can ask for a copy of their data in a usable form. Again, easy if the system knows where each person's data lives and can gather it; painful if that knowledge exists only in the heads of the people who built it. The honest way to treat these is as product features with real engineering behind them — an erasure that truly reaches everywhere, an export that is complete and correct. Build for them early and they are routine, a background job that runs and reports. Discover them late and each request is a manual scramble that pulls engineers off everything else.

Consent and lawful basis, in plain terms

You cannot process personal data just because it is convenient — you need a lawful basis, which is simply a legitimate reason the regulation recognizes. Consent is the one people know, but it is not the only one and often not the best fit. Performing a contract someone entered into, meeting a legal obligation, and certain legitimate interests are all bases too, and choosing the right one matters because it shapes what you must build.

Where consent is the basis, it has to be a real choice — freely given, specific, and as easy to withdraw as to give. A pre-ticked box is not consent, and a withdrawal that the system cannot actually act on is worse than none. In engineering terms, if you rely on consent you must record what someone agreed to and when, and be able to stop the processing when they change their mind. That is a small feature if planned and an awkward one if remembered late. The exact basis for each use is a question for your legal advisor; building the system so it can honor whichever basis applies is our job, and the two conversations are better held at the same time than months apart.

Start with an assessment

The thread through all of this is timing. Every one of these — minimization, retention, access control, residency, erasure, export, consent — is inexpensive when it shapes the design and expensive when it is forced onto a finished system. Retrofitting means unpicking a data model that assumed it could keep everything forever, adding deletion paths through code that never contemplated deletion, and discovering that personal data has quietly spread into places no one mapped. That is why the most useful first step is a clear-eyed look at how data flows through the system you have or plan to build — what you collect, why, where it lives, who touches it, and how erasure and export would actually work. That picture tells you where you are exposed and what to fix, in what order, before it becomes expensive. We do the engineering side; the legal reading you should get from a qualified advisor, and the two fit together well.

Building software that handles personal data?

We run a fixed-fee data-protection design review that maps how personal data flows through your system and turns it into a costed engineering plan — the legal reading pairs with a qualified advisor.

Book a data-protection review