Whitworth's Thread

Gartner thinks 40 percent of agentic AI projects will be canceled by 2027, and DORA found that AI helps enormously or barely at all. In the latest example of everything old is new again, we had the same problem a century ago around reusable parts. We should solve it the same way.

Whitworth's Thread

In 2025, Gartner issued a poll that said it expects more than 40 percent of agentic AI projects to be canceled by the end of 2027. This year, the firm sharpened the claim: by 2027, roughly 40 percent of enterprises will demote or decommission autonomous agents, and the driver it names is governance gaps that only surfaced after a production incident. Note the timing: these gaps surfaced AFTER it rolled out to production. I think it's a pretty safe summary to say that the industry thinks this is a model problem, that the models aren't good enough yet, and that another turn of the capability crank will fix everything. I think this reading is wrong, and this survey points to the reason why.

In a separate survey, DORA's most recent State of DevOps research, drawing on nearly 5,000 technology professionals, found that AI adoption acts as an amplifier: strong positive effects on organizational performance when the underlying platform is strong, and effects near zero when it is weak. In the same dataset, individual productivity rose while delivery throughput and stability declined, with the gains being absorbed by code review, testing, and security sign-off. In the study, they coined the term "downstream disorder," which is kind of adorable, like something a Victorian doctor would diagnose an aristocrat with. As in the Gartner study, none of it is a statement about model quality; it's almost entirely about whether the work was ever specified with meaningful outcomes.

Since the dawn of bureaucracy and organizations, people have been trying to measure outcomes. This is much harder than you might imagine! The biggest problem is that standardization inevitably curdles into bureaucracy. I have personally seen change advisory boards whose actual function was to spread blame too thin to land on anyone, and I have filled in templates whose only reader, ever, was the template's author. Teams route around dead processes within a week, and then there are two processes: the written one, audited and fictional, and the real one, undocumented, resident in the heads of four people. Arguably, that is worse than no standard at all, because now you cannot even see your own variance (I will concede that SOMETIMES it helps, because at least it forces the form-filler-outer to crystallize what they're thinking, but it remains in their head, which is not ideal).

So the distinction I care about is between prohibited variation and controlled variation. A standard that documents its own escape hatches- the four known conditions for leaving the path and who gets told when you do- can absorb surprise. A standard with no exceptions is a lie, and everybody involved knows. Let's figure out how to move forward better.

He Built The Ruler First

In 1841, Joseph Whitworth read a paper to the Institution of Civil Engineers proposing a uniform screw thread for British industry: a 55-degree included angle, with specified radii at the root and crest so that the thread would not concentrate stress where it was most likely to fail. It became the world's first national screw thread standard. Before it, every manufacturer in Britain cut threads to its own proportion; a bolt and its nut were a matched pair fitted to each other by a particular workman, and if you lost the nut, you did not go and get another nut; you went and got a fitter. The year before the thread paper, in 1840, Whitworth had developed what he called end measurements, a technique using a precision flat plane and a measuring screw of his own construction. He had worked it down to a claimed precision of one millionth of an inch, which he later showed off to the public at the Great Exhibition of 1851. (Remember when we had amazing things like that at the World Fair?! I was always so inspired by that. I wish we would bring stuff like that back.)

Metrology first, standard second. He built the ruler before he proposed the rule, and the former's existence made the latter possible.

Look For The Marks

In January 1801, Eli Whitney traveled to Washington and demonstrated interchangeable musket locks before President Adams, President-elect Jefferson, and a room of officials. Ten locks were disassembled, their parts mixed, and then reassembled. The demonstration was a sensation; it secured his contract and made him a fixture in every American textbook as the father of interchangeable parts. In something that will sound really familiar to all the fake-it-before-you-make-it startup people out there, the demo was rigged.

Whitney had marked the parts beforehand so they could be matched back up, and later examination of surviving Whitney muskets showed the components were not interchangeable in any strict sense. Hand filing was still required to make anything fit, and the muskets carry special engraved marks on their parts whose only purpose is to tell an armorer which part belongs to which gun. The man who ACTUALLY achieved interoperability was Honoré Blanc, in France, more than a decade earlier, using jigs, gauges, and master models to hold musket parts to identical tolerances by hand. Jefferson saw Blanc's workshop while serving as ambassador in Paris and wrote home about it, so the American government learned about the real version first and bought the staged one anyway.

The American who finally did it worked at Harpers Ferry, and almost nobody remembers his name. John Hall signed a contract with the War Department in 1819 to produce his breech-loading rifle at a small rifle works on an island in the Shenandoah, and he spent the first several years of it building machines and gauges instead of guns. He built dozens of milling and gauging machines, worked out fixtures that held each part in the same position for every operation, and ran everything against master gauges at every stage of production. In 1826, the government sent an inspection board, which took a hundred of Hall's rifles, stripped them, scrambled the parts in boxes, reassembled rifles from the mixture, and watched the reassembled guns fire. The board reported that the parts could be exchanged with a level of flexibility never before achieved. INTERESTINGLY, most of the contract money had gone into tooling rather than rifles, and Hall's cost per gun came out higher than the ordinary muskets the armories were already producing.

Flash forward to today, every agent demo you have been shown in the last eighteen months is the Whitney demo. I do not mean that as an accusation of fraud, because Whitney sincerely believed he was three years away from the real thing. I mean, the demo works because a human pre-fitted the parts, and you can tell by looking for the marks. Usually, someone quietly reviews the output before it goes anywhere, or an eval set is hand-curated, or one customer is somehow always the reference customer, or a workflow step is described as "and then it just gets handed to the ops team." SWEs and SREs wince at this last phrase, since a genuinely standardized process for rolling out would never involve the word "just." It's a discipline all its own, and it's not something to be papered over.

Sixty-Four Paths, Six Of Them Written Down

In contrast, a confabulation is a silent wrong answer surfacing in a quarterly review nine weeks later. No amount of model tweaking/improvements is going to fix this; it is a property of delegation, and it's been true of every new hire you've ever onboarded. Except now your new hires are agents.

I keep watching teams ask for agentic workflows while their actual onboarding process is forty Slack messages plus a guy named Jerry who knows which of the three staging databases is the real one. Jerry appears in no runbook and no architecture diagram; he is a single point of failure with a mortgage, and he is taking two weeks in August. Four if he's in Europe.

So the practical move is unglamorous, but (nearly) always beneficial. Write the runbook you would hand to a competent new hire, so they can execute it without asking anybody anything. The places where you stall, or where you have to say "and then you kind of know from context," are the places where no agent will operate reliably, and now it falls to you to document. I have written before about the gap between agent ambition and agent infrastructure, about how a loop is only as good as the metric you close it against, and about who is accountable when the thing acts, and they are all this same problem from different sides: you are handing work to a system that cannot ask a clarifying question in the hallway (or Slack), and your operating model was quietly built on the assumption that everything could.

Whitworth's real contribution was never the 55-degree angle; the number was arbitrary, and the Americans later went with 60 and did fine. His contribution was that a bolt made in Manchester was threaded into a nut made in Glasgow by a stranger, which meant you could finally build a machine to make the bolt instead of hiring a man to fit it. Get the measurement, then get the agreement. The thread was never the invention.


Want the boring parts of your data operations to be repeatable enough that something other than a person can run them? Check out Expanso. Or don't. Who am I to tell you what to do?

NOTE: I'm currently writing a book based on what I have seen about the real-world challenges of data preparation for machine learning, focusing on operational, compliance, and cost. I'd love to hear your thoughts!