AI Regulation Is Watching the Datacenter. The Compute Has Already Left the Building.
Sergio Martin Diaz
sergio.md
Governments designed AI oversight around giant, visible clusters of chips. A new paper suggests frontier-scale training could be broken into smaller pieces, scattered across ordinary networks, and hidden in plain sight.
Imagine a government bans the construction of nuclear weapons.
It monitors uranium facilities, inspects enrichment plants, tracks suspicious electricity consumption, and watches large industrial sites from space.
Then someone discovers that the relevant process can be divided across hundreds of smaller locations, none of which looks particularly interesting on its own.
The nuclear analogy is imperfect. AI models are not bombs, GPUs are not uranium, and comparing every technical problem to nuclear proliferation is one of the policy world’s less charming habits.
But the structural problem is real.
A large part of AI governance is being designed around one comfortable assumption: training a frontier model requires an enormous, centralized and highly visible datacenter.
A new research paper asks what happens when that assumption stops being true.
The answer is not reassuring.
The Regulation Watches the Building
Frontier AI models require an absurd amount of computation.
Today, companies usually concentrate that computation inside giant clusters where thousands of specialized chips can communicate through extremely fast connections. The facilities consume remarkable quantities of electricity, require industrial cooling systems and are difficult to conceal.
This visibility has become one of the foundations of “compute governance.”
Instead of trying to inspect every line of code or determine whether a model has developed dangerous capabilities after the fact, regulators can focus on a resource that is easier to count: chips and computation.
The European Union’s AI Act, for example, presumes that general-purpose models trained above 10^25 floating-point operations may pose systemic risk. California’s SB 53 defines frontier models using a threshold above 10^26 operations. Other proposals would require governments to register and monitor clusters above a certain size.
The appeal is obvious.
Software is slippery. Hardware has an address.
Or at least it used to.
The Datacenter Becomes a Network
Robi Rahman, a researcher at the Machine Intelligence Research Institute, modeled whether a sophisticated developer could train large AI models without assembling one giant computing facility.
Instead, the developer would divide the work among many smaller clusters in different locations and connect them through ordinary internet infrastructure.
This would traditionally have been painfully inefficient.
Training requires processors to exchange enormous quantities of information. Modern AI datacenters use specialized interconnects capable of moving hundreds of gigabits per second. A residential or small-business internet connection is not supposed to compete with that.
Recent distributed-training techniques change the equation.
Algorithms in the DiLoCo family allow separate groups of processors to work independently for longer periods before synchronizing. They can also compress the information exchanged between locations. Instead of every processor constantly discussing every microscopic adjustment with every other processor, the nodes work separately and compare notes later.
Think less open-plan office, more competent remote team.
Rahman’s model assumes difficult conditions: connections around 100 Mbps, substantial latency, and individual computing nodes kept below one proposed monitoring threshold. Even under those constraints, the simulations suggest that distributed systems could exceed several of the compute thresholds used in existing laws and governance proposals.
The important discovery is not that distributed training is cheaper.
It is not.
The important discovery is that it may be cheap enough.
The Price of Becoming Somebody Else’s
Enforcement Problem
According to the paper’s estimates, a distributed operation could exceed a proposed international training limit of 10^24 operations with approximately $1.6 million in hardware.
- Crossing the EU AI Act’s 10^25-operation threshold could require about $30.7 million.
- Reaching the estimated raw training compute of GPT-4 would require a more complex setup costing approximately $91.3 million.
- At the extreme end, exceeding California SB 53’s 10^26-operation threshold would require around $3.8 billion.
That last number is not pocket change. It is not even SoftBank pocket change without a presentation deck.
But the relevant question is not whether a garage startup could do it.
The relevant question is whether a large company, a wealthy state or a well-funded organization could find the price acceptable if centralized training were prohibited, sanctioned or closely monitored.
Under those circumstances, inefficiency becomes a fee for privacy.
Thirty million dollars is expensive compared with complying with a normal regulation. It is not expensive compared with acquiring a strategically valuable AI capability that the regulation was specifically written to prevent.
Even billions become plausible when the buyer is a government.
Policy often fails because it compares the cost of evasion with the cost of legal operation. The evader compares it with the value of getting away with it.
Different spreadsheet.
Not Invisible.
Just Invisible to the Tools We Chose.
The most tempting version of this story is that someone could secretly train GPT-4 from a thousand suburban basements while regulators stare helplessly at an empty field.
That is good content. It is also an exaggeration.
A distributed training operation would still leave evidence.
Hundreds of nodes require thousands of chips. The hardware must be purchased, shipped, installed, powered, maintained and eventually replaced. Facilities need staff. Staff need salaries. Companies need suppliers. Suppliers produce invoices. People complain. People quit. People occasionally decide that a government reward is more attractive than their loyalty to an employer with a complicated interpretation of international law.
Rahman’s argument is not that distributed training becomes perfectly invisible.
It is that the detection methods built around centralized datacenters become insufficient.
- Satellite imagery may identify a massive new facility. It is less helpful when the hardware is divided among ordinary industrial buildings.
- Power-grid monitoring may notice a cluster consuming hundreds of megawatts. It is less useful when the load is broken into smaller chunks below the reporting threshold.
- Internet monitoring sounds promising until you remember that the traffic can be encrypted, routed through intermediaries and shaped to resemble less suspicious activity. The modeled systems may need less bandwidth than an average household connection in the United States.
The operation remains detectable.
Just not from one convenient dashboard.
The Great Threshold Game
Regulators love thresholds because thresholds turn difficult judgments into numbers.
- Above this amount of compute, report the training run.
- Above this cluster size, register the hardware.
- Above this revenue level, follow additional rules.
Every threshold also creates a game.
Once companies know where the line is, they can redesign their behavior to remain one millimeter below it. Banks call this structuring. Tax authorities see it constantly. Platforms see it when sellers divide businesses among multiple accounts. Anyone who has ever dealt with corporate procurement knows that humans possess a remarkable ability to turn one purchase into seventeen unrelated invoices.
AI compute is no different.
A rule covering clusters with the computational power of more than 16 H100 GPUs does not necessarily stop someone from using many clusters with the power of 15.9 H100s.
The regulation sees a collection of modest installations.
The training algorithm sees one machine.
This is the central policy problem exposed by the paper. Governance based exclusively on the size of individual clusters may regulate the shape of the infrastructure rather than the actual amount of computation being performed.
The law says, “Do not build a large pile.”
The developer replies, “Of course not. We built 600 small piles.”
Legal departments have charged considerably more for less creative advice.
Memory Is Compute’s Less Famous
Accomplice
One of the paper’s more practical recommendations is that regulators should not measure clusters only by processing speed.
They should also measure memory.
Some older AI chips have relatively large amounts of high-bandwidth memory compared with their computational performance. That makes them useful for evasion: a node can remain below a compute threshold while holding a much larger model locally.
More memory means less information has to move between locations. Less movement means fewer communication bottlenecks. The distributed system becomes more efficient without appearing more powerful under a compute-only rule.
Rahman proposes that cluster-registration requirements should apply when hardware exceeds either a compute threshold or a memory threshold.
This does not close the loophole completely.
It makes the loophole more expensive.
And making evasion expensive matters because it forces an operator to use more nodes, hire more people and generate more financial activity. Each inefficiency becomes another opportunity for discovery.
Good enforcement does not need to make misconduct physically impossible.
It needs to make it operationally ridiculous.
The Most Effective Detection System
May Be an Employee
The paper also recommends something less glamorous than cryptographic chip controls or satellite surveillance: whistleblower programs.
This makes sense for a simple reason.
Centralized infrastructure can be operated by a relatively small team in a secure location. A distributed operation involving hundreds or thousands of sites creates an administrative migraine with a catastrophic number of witnesses.
- Someone has to negotiate the leases.
- Someone has to receive the hardware.
- Someone has to explain why the electricity bill looks like a crypto mine with a PhD.
- Someone has to replace failed equipment in warehouse number 317.
Every additional employee, contractor and supplier expands the human attack surface.
A program modeled on the SEC’s whistleblower system could offer financial rewards and legal protections to people who report unauthorized training operations. It would work with existing hardware, unlike governance mechanisms that must be designed into future chips.
Not every policy problem requires a new blockchain.
Sometimes you pay the operations manager.
AI Governance Cannot Be a Datacenter
Zoning Code
The paper does not prove that a secret organization could effortlessly reproduce the best commercial AI systems tomorrow.
Its highest-scale results involve significant extrapolation from distributed-training experiments conducted at smaller scales. Training across hundreds or thousands of locations would remain technically difficult. Hardware failures, data quality, software coordination and model efficiency all introduce uncertainty.
But policymakers do not get to ignore a vulnerability until somebody publishes a polished case study titled How We Broke the Treaty.
The useful lesson is broader.
AI governance is currently attracted to compute because compute appears physical, concentrated and measurable. Those properties make it easier to regulate than algorithms or capabilities.
Distributed training weakens all three.
The response cannot be to abandon compute governance. Chips are still more trackable than code, and a distributed operation may actually need more of them than a centralized one.
The response is to stop pretending that one metric and one monitoring system will be enough.
A credible framework would combine chip registries, compute and memory thresholds, financial audits, supply-chain intelligence, surprise inspections and rewards for insiders. No individual measure solves the problem. Together, they make concealment expensive and discovery more likely.
That is how serious regulation usually works.
Not with one perfect rule, but with several mildly annoying systems that become deeply unpleasant when encountered simultaneously.
The Laws Are Young.
The Workarounds Are Not Waiting.
AI governance is still being written.
That is an advantage. The rules have not yet calcified into ceremonial compliance documents maintained by seventeen committees and understood by nobody.
It is also a risk.
The industry is changing faster than the assumptions behind the legislation. A law designed around the infrastructure of 2024 may encounter the training techniques of 2028 and discover that it has been regulating a building that no longer needs to exist.
The lesson from distributed training is not that AI regulation is hopeless.
It is that regulators cannot govern a moving technology using a still image.
The datacenter was a convenient chokepoint.
Convenient chokepoints have a habit of becoming optional.