Acceptable for Inference

Google tested its TPUs for orbit and the memory came back with uncorrectable bit flips at a rate it called likely acceptable for inference. Read that phrase again. We are about to make correctness a per-workload dial set partly by radiation, and the model will not tell you which setting it used.

Acceptable for Inference

Starcloud closed $170 million at a $1.1 billion valuation this month, the fastest company in Y Combinator's history to reach a billion. Starcloud-2 launches later this year carrying Blackwell B200s to run commercial cloud workloads in orbit for customers that already include AWS and Google Cloud. This is pretty rare for a YC company! ACTUAL paying tenants, this year, in a rack that is going around the Earth every ninety minutes.

When Google published Project Suncatcher, the press took the obvious angle: Google wants data centers in space, fleets of TPUs linked by free-space optics into kilometer-wide arrays of 81, two test birds going up with Planet by early 2027. Solar power that never sets, which seems exactly right! Let's do it!

But, Google ran its TPUs through a particle accelerator to simulate the dose of low-earth orbit, and the compute chips came through fine. The high-bandwidth memory took uncorrectable errors that the error-correcting code could not catch and repair, at a rate Google described as "likely acceptable for inference."

Not "acceptable" for anything, but "acceptable for inference". This is pretty specific guidance that running in a space is only suitable for a specific job it has in mind will forgive the occasional wrong bit.

This makes sense! A model writing the seventh paragraph of a product description genuinely does not care if one weight in one layer got nudged by a passing cosmic ray. The output was a probability distribution to begin with, so a little noise in the machine is a rounding error inside a process that was already rolling dice.

Precision was always a marketing choice

We have spent the entire AI era letting people believe these systems are precise, and orbit just makes the imprecision physical. Down here, the fuzziness hides inside phrasing that sounds authoritative. Up there, it's a photon flipping a one to a zero in a memory cell, and the model downstream will report the result with exactly the same confidence it would have had if the bit were correct. I wrote a while back about a support chatbot that told a customer they had 365 days to return a product when the real policy was 30, every dashboard green, the model perfectly poised while it was flatly wrong. Now picture that same confidence, except this time the error was injected by the sky.

For a chatbot, who cares. The trouble starts the instant somebody wires a forgiving workload to an unforgiving job. Starcloud has filed to put 88,000 satellites in orbit to process data rather than relay it, and somewhere in the addressable market for eighty-eight thousand orbiting accelerators is a company that will run something that counts on hardware whose spec sheet says, in so many words, good enough to be wrong sometimes. The problem isn't JUST that it can be wrong, but that it can be wrong silently, since nobody in that chain is going to be told which rack the answer came from. That is the entire product promise of cloud: you don't think about the hardware. It is a very good promise right up until the hardware develops opinions.

The rot is already in the building

If you're about to file this under "space is weird," don't. Silent data corruption is not JUST an orbital phenomenon. Meta went looking on its own fleet and found corrupted computations coming out of perfectly healthy-looking CPUs at a rate high enough to matter across a datacenter, and the industry now has an Open Compute working group and a whitepaper about it specifically because inference multiplies the blast radius: one marginal device quietly wrong, hundreds of thousands of inferences an hour, every one of them delivered to a customer with full confidence. The causes are mundane and unfixable, timing violations, aging, marginal defects, temperature, voltage, and yes, cosmic rays hitting silicon at sea level.

So orbit is not introducing a new bargain; it is turning up the gain on one we already made and mostly declined to discuss. What Google did that's genuinely new is write the terms down. Every terrestrial datacenter is running some rate of silent corruption it does not advertise and cannot fully measure. Google put a number next to it, attached a workload class, and called it acceptable. That candor is rare enough here that it reads as alarming, but it shouldn't. That candor should come stapled to the side of every output, like an FDA label.

We named this bargain once already

Distributed systems made exactly this trade a long time ago, on the ground, and we even gave it a name. Eventual consistency. We decided that for a shopping cart or a like counter, the right answer soon-ish beats the exact answer slower, and we built half the modern internet on top of that call. It sits close enough to the point of this whole blog that it's in the subtitle. But the expensive lesson, the one every architect learns once and never forgets, was working out which systems you are absolutely not allowed to make eventually consistent. Amazon runs its catalog eventually consistent and its payments emphatically not, and the entire art was knowing exactly where that line sat.

The orbital memory result is that same fork, pushed down to the level of a single bit and handed to radiation to decide. Correctness is about to become a per-workload dial that gets set, in part, by how much cosmic radiation a given satellite happened to eat that week. And the thing generating the answer is not going to print which setting it was running on.

So the question worth asking about compute in space was never whether we can do it. We can (at least to some degree), the first paying workloads go up this year, and the solar-power argument is real and worth taking seriously. The question is who keeps track of which answers came back from a place where the memory does not reliably hold, because the model certainly won't volunteer it. It'll sound exactly as confident either way. It always does.

Google, to its credit, told us the setting. Ask your own vendors what theirs is.


Want to learn how intelligent data pipelines can reduce your AI costs? Check out Expanso. Or don't. Who am I to tell you what to do.

NOTE: I'm currently writing a book based on what I have seen about the real-world challenges of data preparation for machine learning, focusing on operational, compliance, and cost. I'd love to hear your thoughts!