AI Data Centers 101: What a 100K+ GPU Cluster Actually Needs
A 100,000-GPU AI training cluster is as much a power station and a cooling plant as a compute installation. This plain-language tour covers the four hard constraints — power, cooling, networking fabric, and storage — and why the physical layer is now the binding limit on AI progress.
Power: the first wall you hit
A 100,000-GPU cluster does not draw power the way a data center draws power. The headline number is usually in the 100–150 MW range for the IT load alone, before you add the substations, switchgear, and backup generation that a site of that size requires. That is not a number you can satisfy with a upgraded commercial lease and a couple of feeders from the local utility. It is a number that gets utility planners, city councils, and grid operators involved before the first GPU is racked.
The 100,000-GPU cluster is, in practical terms, a small power station with a few ten thousand silicon chips attached to it. The compute story — model sizes, token budgets, training runs — is the part that makes the headlines. The power story is the part that decides whether the thing gets built at all.
Cooling: air stops working at this scale
Once you pass a certain rack density, air cooling stops being a viable strategy. The 100,000-GPU cluster lives in a regime where the heat is too concentrated for room-level air to carry it away economically. That is why liquid cooling — cold plates on the GPUs, and in some designs immersion cooling for the whole rack — has moved from a niche option to the default assumption at this scale.
The cooling plant is not a side system. It is a first-order part of the facility design, competing for space, power, and capital with the IT equipment itself. The choice of cooling topology affects the building, the power distribution, the operations team, and the reliability story. A 100,000-GPU cluster without a credible cooling plan is not a cluster — it is a proposal.
Networking fabric: the glue that holds the 100,000-GPU cluster together
A 100,000-GPU cluster is only as useful as its interconnect. Training runs at this scale are bounded as much by how fast the GPUs can talk to each other as by how fast any one GPU can do math. The networking fabric — spine-and-leaf, optical switches, the topology that decides which GPU can talk to which other GPU and how cheaply — is what turns 100,000 individual accelerators into a single machine.
This is the layer that gets less attention than the chips and the power, but it is the layer that decides whether the 100,000-GPU cluster trains a model in a week or a month. The fabric is also where a lot of the physical-layer innovation is happening: optical interconnects, co-packaged optics, and new switch architectures are all fighting to reduce the communication penalty that dominates large-scale training time.
Storage: keeping 100,000 GPUs fed
A 100,000-GPU cluster that spends a meaningful fraction of its time waiting for data is a 100,000-GPU cluster that is not earning its keep. The storage pipeline — parallel filesystems, caching layers, the data-loading architecture that gets training tokens from disk to GPU memory fast enough to keep the compute busy — has to be designed for the same scale as the rest of the facility.
At this scale, the storage story is not about capacity. Tape archives and object stores handle capacity cheaply. The storage story at the 100,000-GPU level is about throughput and latency: getting the right bytes to the right GPU at the right time, across a facility that spans multiple buildings, multiple power domains, and a networking fabric that is already doing its best to stay out of the way.
Why the physical layer is the binding constraint
The software stack — frameworks, compilers, distributed training algorithms — keeps improving. The chips keep getting faster. But the 100,000-GPU cluster is constrained by things that do not improve on a software roadmap: the available power on the grid, the water or air that carries heat away, the fibers and switches that link the racks, the disks and caches that feed them.
This is why the AI data center story has shifted from a compute story to a facilities story. The 100,000-GPU cluster is real, and it is being built. But every one of them is, underneath the marketing language, a negotiation with physics — with how much power a site can deliver, how much heat a cooling plant can carry, how fast a fabric can move tensors, how quickly storage can feed a hungry cluster.
The binding constraint on AI progress is no longer just how many parameters we can train. It is how many 100,000-GPU clusters we can physically build, power, cool, wire, and feed — and how fast.