# Privacy by architecture: the data you never collect can't leak

- Published: 2026-09-30
- Reading time: 8 min
- Tags: Privacy, Data, Architecture

The best privacy policy isn't code that deletes data — it's a schema that never made room for it. Data minimization is an architectural decision, not a legal clause.

### Minimization: don't collect it, don't keep it

The most common mistake in data design is assuming "having it doesn't hurt." Every field a schema adds carries a hidden cost: a surface for leakage that didn't exist before, a column that eventually shows up in a backup dump or a misconfigured log, one more thing that has to be justified against a deletion request, an audit, or a breach. The practical rule is simple: if a feature doesn't need a field today, that field shouldn't exist — "we might need it later" is not a justification. "Might need it later" makes today's cost real for a benefit that is still hypothetical. If it turns out you really do need it later, adding a column to an existing schema is a migration; unwinding a years-old habit of collecting data you never needed is a project. Retention follows the same logic. Data that gets collected but never deleted past the point it was defined for carries exactly the risk of data that was never collected, except it actually exists.

### Design the schema around the feature's need, not "might be useful"

The right way to design a schema starts from the feature, not from the data. "What data could we collect" is the wrong question. The right one is: "what does this specific feature need to actually work?" If a feature only needs to know whether a user belongs to a group, it doesn't need to store the full interaction history of that group. This discipline gets harder when a product or analytics team wants to "keep everything, we might want to analyze it later." That request deserves to be confronted with its real cost, not automatically approved or automatically rejected.

### Retention as a first-class field

Any table holding personal or sensitive data should have a retention policy written into the initial design, not bolted on later as a separate "compliance" project. That means an expiry field or a scheduled cleanup job from day one of the schema — not something written in a hurry the day the first deletion request arrives. Retention designed in from the start is cheap. Retention added afterward is expensive, because you have to work out which old records belong to which user without breaking other relationships that are supposed to remain.

### The honest cost to analytics

Data minimization isn't free. A product team will occasionally be unable to answer a future question because the data it needed was never collected. That's a real cost, and it should be stated honestly rather than hidden behind the slogan "privacy matters." The right move is to put that cost on the table at the moment the decision is made about what to collect — not to claim there is no cost at all. Sometimes the right answer is to keep aggregated, anonymized data while discarding the individual-level record. That's a deliberate trade-off to choose, not something that happens by default because nobody took the time to design it.

### Communication and news products need extra care

A product that carries people's conversations and a product that attaches a trust or source status to news content are both examples of data more sensitive than an ordinary signup form. A communication product generates metadata — who spoke to whom, and when — that is revealing even without the content itself. A news product that labels content with a trust or source signal is, in effect, also generating data about a user's beliefs and preferences, not just their behavior. In both cases the rule above holds, just with a lower tolerance for error: if you're not sure a field is needed, drop it. If you're not sure how long to keep something, pick the shortest retention that still works, not the longest one that might someday be useful.

### The logistical blind spots: backups, logs and caches

A minimization policy that only applies to primary tables and not to secondary copies of the same data is effectively half a policy. A full database backup holds the same sensitive data with the same sensitivity, but usually with weaker access control — backups get audited less and live longer. If a retention policy only touches the live table while old backups still hold the deleted record, the user's deletion was a half-kept promise. The same story applies to application logs. A debug log line that thoughtlessly prints a full request payload can copy exactly the field you removed from the database into a log file that gets retained for a year. Caches carry the same risk: sensitive data sitting in a cache layer with its own expiry policy that nobody separately wrote a privacy rule for. Real minimization means asking this question at every layer data passes through, not only where it's stored permanently.

### When full deletion isn't possible: anonymize instead of delete

Sometimes deleting a record outright conflicts with another requirement — a financial transaction with a mandated retention period, or a record that's a foreign key for other rows. In those cases, the right answer is usually not immediate deletion but anonymization: replace identifying fields with meaningless values while keeping structural relationships intact. That's different from just flipping a "deleted" flag and leaving everything else untouched — the latter isn't deletion, it's concealment. The design should specify up front which fields actually get erased on a deletion request and which only get anonymized, and that decision should be documented in the schema — not something a new support team member has to rediscover from scratch every time.

### Shrink the blast radius of access, not just the data stored

Minimization isn't only about what gets collected; it's also about who can reach it, and under what conditions. A field that genuinely needs to be stored doesn't automatically mean every internal service should be able to read it directly. Broad access to sensitive data is itself a form of hidden collection: the more places a piece of data is visible from, the larger the leak surface, even if the stored data itself is minimal. Good design solves this with access layering: a service that only needs to know whether a user belongs to a group shouldn't have a direct path to the table holding that user's full profile — it should go through a narrow interface that answers exactly that one question. That layering has an engineering cost, and it's often skipped with the excuse that "it's all internal, we trust each other" — the same excuse that, every time, turns an internal leak into an external one.

### A short checklist

What specific feature does this field enable today? Where is this table's retention policy written down — in the schema, or only in someone's head? If a full deletion request for a user landed today, how many tables would need manual review? Has anyone stated the analytics cost of this decision out loud, or is it just assumed to be zero?

### This conversation belongs with the client too, not only inside engineering

Most requests for extra data collection don't originate inside the engineering team; they come from a product or business decision whose privacy cost hasn't been weighed yet. When a client says "store this field too, we might need it later," the most honest engineering response isn't to accept it unconditionally or reject it unconditionally — it's to put the cost and risk on the table at that same moment: how sensitive is this field, who ends up with access to it, and how much extra work does it create if a deletion request arrives someday. That's what clarity about scope and risk actually means in practice — not just stating how long something takes and what it costs, but stating what responsibility a field adds to the system down the line, before it lands in the schema, not after.


---

Source: https://larsima.com/fr/insights/privacy-by-architecture
Organisation: LARSIMA (شرکت فن‌آوران توسعه لار سیما), registered in Iran, no. 318515, since 2007-12-02.
