July 20, 2026
Modular Data Center for AI: NVL72 Density, Speed, and the Retrofit Trap
A GB200 NVL72 rack draws ~120 kW; legacy halls cap near 30 kW. Modular data centers for AI skip the retrofit and deploy in months, not years.

A modular data center for AI is a factory-built facility engineered from the start for high-density, liquid-cooled GPU racks, the profile that legacy data centers can't host without near-total reconstruction. An NVIDIA GB200 NVL72 rack draws roughly 120 kW, with deployed systems reported at 130–132 kW under full load (NVIDIA / HPE, 2025), while conventional air-cooled halls top out around 20–30 kW per rack (Introl, 2025). Above that ceiling, retrofitting means new power distribution, new liquid-cooling plumbing, and reinforced floors: "reconstruction, not incremental retrofit" (Global Data Center Hub, 2026). Purpose-built modular sidesteps the trap by treating density and liquid cooling as design premises, not upgrades.
This post covers the rack-density explosion, why legacy retrofits fail the math, how modular avoids the problem, the difference between inference and training infrastructure, and what idle GPUs cost while you wait.
Why has AI broken the rack-power assumption?
Because a single rack now draws what a room used to. That one fact invalidates most of the installed base.
The NVIDIA GB200 NVL72 packs 72 Blackwell GPUs and 36 Grace CPUs into one rack-scale system and runs at about 120 kW nominal, with real deployments at 130–132 kW under full load (NVIDIA / HPE, 2025). For context, a single H100 draws around 700 W, and an 8–16 GPU training rack reaches 40–80 kW (Global Data Center Hub, 2026). Legacy colocation from the 2000s was built for 3–5 kW per rack; cloud-era hyperscale for 8–15 kW.
Air cooling can't follow. Conventional raised-floor air tops out around 20–30 kW per rack, and even purpose-engineered air designs rarely clear 30–40 kW before the required airflow exceeds what the floor can physically move (Introl, 2025). NVIDIA doesn't leave it ambiguous, the NVL72 mandates direct-to-chip liquid cooling, with a 20–25°C inlet and roughly 80 L/min flow (NVIDIA, 2025). The full thermal breakdown, from rear-door exchangers to immersion, is in liquid cooling for data centers, and what Blackwell specifically demands of a facility is in Nvidia Blackwell explained.
Why can't you just retrofit a legacy data center?
Because three independent variables each disqualify most existing buildings, and together they make the ROI collapse.
Power density is the first. AI-ready racks need 30–100 kW-plus, a 6× to 20× jump over legacy 3–15 kW designs (Introl, 2025). Floor loading is the second, a fully equipped liquid-cooled rack can be 3–5× heavier than its air-cooled predecessor, while older raised floors are often rated at only 1,000–1,500 kg/m² (Introl, 2025). Liquid-cooling plumbing is the third: the pipes, manifolds, leak detection, and structural reinforcement simply don't exist in air-cooled facilities, and adding them "amounts to reconstruction, not incremental retrofit" (Global Data Center Hub, 2026).
Global Data Center Hub documented a four-year-old, OCP-grade facility running an efficient PUE of 1.15, designed for 8–12 kW per rack, that was "structurally disqualified" when a customer asked for 40–50 kW. The retrofit ran to nine figures and the ROI didn't work (Global Data Center Hub, 2026). Their diagnostic is blunt: a credible AI facility delivers 40 kW-plus per rack without new capex, while upgraded cloud-era facilities reach only 20–25 kW. The gap is structural. Plenty of "AI-ready" marketing sits on top of legacy engineering, a stranded-asset risk dressed up as an opportunity.
How does modular avoid the retrofit problem?
By building for the workload from inception instead of apologizing to it afterward.
Purpose-built modular AI units integrate liquid cooling, high-voltage power distribution, and reinforced structural loading in the factory. This shifts thermal management from a site-engineered constraint to a scalable capability designed around the workload (Introl, 2025). Where air-cooled modules max out around 30 kW per rack, liquid-cooled modular designs are built for the 40–100 kW densities AI demands. ModulEdge units support 5–150 kW per rack, spanning a modest inference node to a full NVL72-class deployment without re-engineering the building.
Hyperscalers validate the direction. Microsoft, Google, and Amazon are all expanding modular programs, and vendors are shipping prefabricated AI modules claiming up to 60% deployment-time reduction (Introl, 2025; Delta, 2026). If you already have the silicon and are staring at the "now what," that operational gap is exactly what the GPU data center problem nobody budgets for unpacks.
Inference or training, which is this actually for?
Both exist, but they want different infrastructure, and being honest about the difference is where most vendors get sloppy.
Sources: EdgeCore, 2025; Introl, 2025; Dell'Oro, 2026.
Training tolerates milliseconds of external latency, so it chases the cheapest power and land. Think gigawatt-scale campuses. Inference sits in the live user path, so it runs close to users across many regional sites (EdgeCore, 2025). Dell'Oro notes that while much AI capex has funded training, "inferencing will likely become a larger capex driver going forward" as reasoning models spread (Dell'Oro, 2026).
Where ModulEdge fits: distributed, latency-sensitive inference and GPU colocation at 40 kW per rack and above, plus enterprise and private AI. Frontier training at gigawatt scale is industry context, not the pitch, and inference is where the market is heading anyway, so it's the stronger story. The chip-level version of this argument is in GPU vs LPU vs NPU, and the rack-power translation for inference specifically is in edge AI infrastructure for inference.
What do idle GPUs cost while a site gets built?
More than most capex models admit. This is the number that makes time-to-power a financial argument, not a convenience.
Even strong in-house teams need 8–12 months to reach production readiness, during which the GPUs are often already installed and depreciating (Introl, 2025). Deployment delays typically cost 20–30% of cluster capex per half-year; a six-month delay on a 512-H100 cluster represents over $9 million in lost value and unrealized revenue (Rafay, 2026). On a five-year straight-line schedule, six months idle burns roughly 10% of asset value with zero output. Combined depreciation and forgone rental can exceed $100 per H100 per day (Introl, 2025).
Modular answers with speed. Liquid-cooled AI modules have been deployed in 8–10 months, and one operator stood up 36 micro modular data centers across 20 cities in 11 months, three times faster than traditional construction at about 40% lower cost (Introl, 2025). Because idle GPUs bleed money daily, every month a modular deployment saves converts directly into avoided depreciation and captured revenue.
The capex backdrop makes the urgency literal: worldwide data center capex rose 57% in 2025, GPUs and custom accelerators now account for about a third of it, and 2026 spend is forecast past $1 trillion (Dell'Oro, 2026). North American colocation vacancy hit a record-low 1.4% at the end of 2025, and CBRE notes inference AI is "redefining demand toward more regional and distributed data centers" (CBRE, 2026). That's the modular thesis, stated by the market. For the full delivery model, start with the definitive guide to modular data centers.
Frequently asked questions
How much power does an NVIDIA GB200 NVL72 rack use? NVIDIA's nominal spec is about 120 kW per rack, and deployed systems have been reported at 130–132 kW under full load (NVIDIA / HPE, 2025). That is roughly 15–40× the density of a legacy 3–8 kW rack, which is why it breaks conventional data center design assumptions.
Why can't a GB200 NVL72 be cooled with air? Conventional raised-floor air cooling tops out around 20–30 kW per rack, and even purpose-built air designs rarely exceed 30–40 kW (Introl, 2025). At 120 kW the required airflow exceeds what any air system can move, so NVIDIA mandates direct-to-chip liquid cooling with a 20–25°C inlet and roughly 80 L/min flow (NVIDIA, 2025).
Can an existing data center be retrofitted for AI racks? Usually not economically. AI racks require 6–20× more power density, 3–5× heavier floor loading, high-voltage distribution, and liquid-cooling plumbing that "amounts to reconstruction, not incremental retrofit" (Global Data Center Hub, 2026). One documented four-year-old facility faced nine-figure retrofit costs and negative ROI, leaving it structurally disqualified.
How is inference infrastructure different from training infrastructure? Training uses very large, tightly coupled GPU clusters in a few locations and tolerates external latency, so it chases cheap power and land. Inference is latency-sensitive and distributed, running closer to users across many regional sites, often on commodity Ethernet (EdgeCore / Introl, 2025). Inference is expected to become a larger share of AI capex going forward (Dell'Oro, 2026).
What does it cost to leave GPUs idle while a site gets built? A lot. Deployment delays run 20–30% of cluster capex per half-year, and a six-month delay on a 512-H100 cluster is over $9 million in lost value (Rafay, 2026). Combined depreciation and forgone rental revenue can exceed $100 per H100 per day, which is why time-to-power, where modular wins, is the metric that matters (Introl, 2025).
Why is a modular data center better suited to AI than a converted building? Purpose-built modular units integrate liquid cooling, high-voltage power, and reinforced loading from the factory, supporting the 40–100 kW densities AI needs rather than retrofitting a building designed for 8–15 kW (Introl, 2025). They also deploy far faster, liquid-cooled AI modules in 8–10 months versus 18–36 for traditional builds, which directly reduces idle-GPU losses.
