The Drums Out Back

IBM put the average breach at $4.99 million. In January the CNIL fined one French carrier €42 million, and part of the finding was that it kept data it had no business still having. Turns out "storing everything" has a cost that goes beyond hard drives.

The Drums Out Back

IBM published its 2026 Cost of a Data Breach report at the end of July. The global average is now $4.99 million per breach, up 12 percent year over year and a record for the study, which this round covered 602 organizations hit between March 2025 and February 2026. American breaches ran more than double the global figure. Healthcare stayed the most expensive industry to be breached in, at $6.64 million a pop. And one in four malicious breaches is now AI-enabled, averaging around $6 million, which is a 56 percent jump in a single year.

The report prices an event for the first time in a while that I actually like. But almost nobody is reading the other half of the sentence; it's half the event, and half the size of the pile. We need to look at both.

The Cleanest Illustration Available Is French

On January 13, the CNIL fined Free Mobile €27 million and Free €15 million, €42 million between them, over an intrusion in October 2024 that reached personal data on roughly 24 million subscriber contracts, including bank account numbers for customers who held both services. Three findings: the data was not adequately secured, the notification to affected people was inadequate, and the retention was illegal.

The third item is pretty uncommon to say the least. The regulator determined that the carrier was holding millions of records on people who were no longer customers, well past any period it could justify, when what it actually needed to keep for accounting purposes was a much smaller slice. So when the attacker came along, they did not steal what Free Mobile needed to run its business; they stole information that Free Mobile had no lawful reason to still possess, sitting in the same system as the material it did.

The intrusion was the trigger, but the retention was the multiplier. One of those two things was under the company's control for years in advance, at essentially zero cost. Less than zero cost, they could have deleted it and saved money!

Article 5 of the GDPR has said since 2018 that personal data must be adequate, relevant, and limited to what is necessary, and that it must not be kept in identifiable form longer than the purpose requires. Violating the Article 5 principles sits in the top tier of Article 83, which reaches €20 million or four percent of worldwide annual turnover, whichever hurts more. Cumulative GDPR fines across the EEA passed €7.1 billion by the start of this year. Minimization has been law for eight years and mostly treated as a policy document that lives in a wiki nobody reads.

Somebody Already Ran This Experiment With Actual Barrels

In 1920 a company called Hooker Chemical bought a partially dug canal in Niagara Falls, and for the next 33 years it used the thing as a disposal site, putting roughly 22,000 tons of chemical waste into the ground. The logic was completely sound, storage was cheap, the land was theirs, the practice was legal, and a fair amount of what went in the hole was technically feedstock, material with a real industrial value if anyone ever wanted to go get it.

In 1953 Hooker sold the site to the local school board for one dollar. The deed carried what became known as the Hooker clause, a written disclaimer saying the company would not be responsible if anybody got sick or died from the waste buried there. The local board proceeded to build school and a neighborhood went up on top.

Then the 1970s happened, families were relocated, Congress held hearings that are still studied as an oversight case, and on December 11, 1980, Carter signed CERCLA. And the statute did two things that have relevance today.

It made liability strict, meaning you do not get to argue that you followed the industry standard or that you were not negligent. And it made liability retroactive, meaning it reached backward to conduct that was perfectly lawful when it happened. The Hooker clause, a real contract, negotiated and signed by consenting parties, turned out to be worth nothing. Occidental Petroleum, which bought Hooker later, inherited the whole thing.

It's true that retroactive strict liability is a sledgehammer, the transaction costs have been genuinely awful, and a meaningful fraction of the money went to lawyers arguing about allocation rather than to anybody's groundwater. Nobody should hold CERCLA up as model legislation, but it was pretty groundbreaking (pardon the pun) for the time. What it gives us, for our purposes, is a demonstration of what a legislature actually does once a stored-material problem gets bad enough and public enough. It does not grandfather you and it does not care what your contract says.

So when somebody tells me their data retention exposure is handled because the DPA covers it, or because the terms of service disclaim it, or because everything was collected lawfully under the rules that applied at the time, I think about a signed piece of paper from 1953 that did all three of those things.

What Minimization Actually Costs You, And When

The reason "keep everything, decide later" won is that it was the rational choice for about fifteen years. Storage got cheap faster than anyone could develop taste, the analytics you would want were genuinely unknowable in advance, and every data scientist who ever got told "we dropped that column in 2019" learned to hoard. I have been the person arguing for keeping the raw feed (and am currently paying penance). Sometimes that argument is right.

But that's no longer the general case. Heck, in many ways it's the exception.

A record is created at some edge of your system. A meter, a handset, a claims form, a badge reader, a log line. At that instant, exactly one copy exists, in one jurisdiction, under one owner, and dropping a field costs you one line of configuration. That is the cheapest that decision will ever be, by orders of magnitude, and it is the only moment where "delete" means the thing deleted is gone.

Now let it move. It lands in object storage, gets picked up into a warehouse, gets denormalized into three marts because three teams wanted different grain, gets a feature-store copy for the model, gets a nightly backup with a 90-day cycle, gets replicated to a second region for durability, gets pulled into a vendor's SaaS for enrichment, and gets exported once into a notebook by an analyst who left in March. Call that eight to ten locations, and I am being conservative, because I have not counted the CI fixture somebody generated from prod or the Slack thread with the screenshot.

Every one of those is now a separate deletion obligation, a separate discovery obligation, a separate breach surface, and a separate thing you have to be able to prove about. Not just assert. Prove, to a regulator, with evidence, on a clock. Deleting a field from ten places under audit is not ten times harder than dropping it once at the source. That kind of work gets a program manager, a status color, and a standing Thursday call, and at the end of it the answer is usually "we believe so."

One decision at creation, or ten proofs later, you only get to choose one. And if you don't choose, you default to the ten proofs version. Good luck.

The Part That Should Actually Bother You

We built an entire generation of data infrastructure whose default posture is accumulate now, govern later. And nobody has ever once meant anything by "govern later." It's a sentence you say to end a meeting, and everybody nodding along knows it's bullshit while they're nodding. LATER NEVER COMES. It does not come because there is no forcing function, no deadline, and no individual whose bonus depends on the absence of a thing.

What does eventually arrive is an incident, or a subpoena, or a supervisory authority with a questionnaire, and by then the population of records you are answering for was fixed years ago by a default nobody chose deliberately. It's not that there was a bad call; at least that could be forgivable. That there was no call. A retention policy emerged as a side effect of a storage tier being cheap.

For the regulated and critical-infrastructure crowd, and I spend a lot of my time in rooms with exactly those people, the question in the room has already shifted. It used to be some version of what could we learn from this if we had it all. Increasingly it is can we demonstrate we never held it. Those two questions want opposite architectures, and only one of them is compatible with a pipeline that ships everything to a central lake first and sorts it out downstream. If the filtering, the redaction, the classification, and the retention decision do not happen in the environment where the record was born, they are happening after the liability has already been created and copied.

I have written before that data loses economic value as it sits, and that your catalog is wrong and will keep being wrong. At the end of the day, the value of a record decays, the accuracy of your inventory of it decays, and the liability attached to it does not decay at all. It sits flat, or it goes up when somebody passes a statute. You are holding an asset that amortizes against a liability that does not.

It's ALSO true that SOME of this material genuinely is an asset, and SOME of it is regulatorily required to be kept, and the retention schedule for a medical record is not a thing an engineer gets to have an opinion about. Fine. Good. That is the point. The categories are different, they have different clocks, and a system that cannot tell them apart FROM THE MOMENT OF CREATION has already decided to treat all of it as the most dangerous thing in the pile. Classification at the point of creation is what lets you keep the valuable part at all, which is why I get twitchy when people file it under compliance and hand it to the team with no engineers.

The Clause Was Signed

Time and jurisdiction draw the line between an asset and a liability. And you do not control either one of them. Hooker Chemical had a legal practice, a real commercial rationale, and a signed indemnity. What it did not have was a view into the future, and in 1980 folks decided they were wrong and should have always done better.

You do not get to know today which of your columns becomes a drum out back. What you get to decide is how many copies of it exist by the time somebody asks.


Want to filter, redact, and classify data where it is created, so the copy you keep is the one you meant to keep? Check out Expanso. Or don't. Who am I to tell you what to do.

NOTE: I'm currently writing a book based on what I have seen about the real-world challenges of data preparation for machine learning, focusing on operational, compliance, and cost. I'd love to hear your thoughts!