Blog

/

NVIDIA GB300 NVL72: What One Rack Delivers, and What It Demands

August 9, 2026

NVIDIA GB300 NVL72: What One Rack Delivers, and What It Demands

GB300 NVL72 draws 132-142 kW per rack, holds 20 TB pooled HBM, weighs 1,580 kg. Full specs, capacity math, and what your building must provide.

The NVIDIA GB300 NVL72 is a single rack containing 72 Blackwell Ultra GPUs and 36 Grace CPUs, wired into one coherent NVLink domain that pools 20 TB of HBM3E memory. It draws 132 to 142 kW at nominal load and peaks near 155 kW, weighs 1,580 kg populated, and rejects roughly 90% of its heat into liquid. One rack produced 673,936 tokens per second on DeepSeek-R1 in MLPerf Inference v6.0. The average enterprise rack, according to Uptime Institute's 2025 survey, runs at about 9 kW.

That last comparison is the whole problem. What follows covers what GB300 NVL72 actually is, why NVIDIA stopped selling chips and started selling racks, what the thing physically demands from a building, how many users one rack can really serve, and what it costs to find out.

What GB300 NVL72 actually is

Start with the shape of it. Eighteen compute trays at four GPUs each, nine NVSwitch trays and eight power shelves, all of it inside one 48U cabinet 600 mm wide that sits on a footprint of about 0.64 m².

Inside that cabinet sit 72 Blackwell Ultra GPUs at 288 GB of HBM3E apiece, 36 Grace CPUs, and a passive copper backplane carrying 130 TB/s of aggregate NVLink 5 bandwidth. Every GPU can read every other GPU's memory at that speed across the backplane itself, which is a rack-internal path where you would normally expect a network hop.

SpecificationGB300 NVL72
GPUs72 × Blackwell Ultra (B300)
CPUs36 × Grace (Arm Neoverse V2)
HBM3E per GPU288 GB, 12-Hi stacks, 8 TB/s
Pooled HBM per rack20 TB
NVLink 5 aggregate130 TB/s
NVFP4 dense / sparse1,080 / 1,440 PFLOPS
Rack power, nominal132–142 kW
Rack power, peak (EDPp)~155 kW
Populated weight1,580 kg
Heat to liquid~90%
Facility water classASHRAE W45 (2–50 °C supply)

NVIDIA's own GB300 NVL72 specifications leave out one thing, which is a rack power figure. Every number in this guide's power section comes from OEM documentation instead, and that gap is itself worth noticing.

We've covered the Blackwell architecture in detail, including what Blackwell demands from a facility. This guide starts one level up, at the rack, treated as a unit of purchase, because that is now what it is.

B300, HGX B300, GB300, NVL72: which one you're actually buying

Four names, and vendors use them loosely, even though they are not interchangeable and picking the wrong one is a seven-figure mistake.

B300 is the GPU. Blackwell Ultra silicon carrying 288 GB of HBM3E at 1,400 W, which you never buy on its own, since it arrives inside something.

HGX B300 is an eight-GPU baseboard with an x86 host, carrying roughly 2.1 to 2.3 TB of HBM per node depending on whose datasheet you read, air-coolable in some configurations, and it drops into a conventional server chassis. Its NVLink domain stops at eight GPUs. Most enterprises should be looking at this one (almost nobody tells them so).

GB300 is the superchip, Grace CPUs coupled to Blackwell Ultra GPUs over NVLink-C2C on an integrated Arm platform. The CPU comes with it, which rules out bolting a GPU onto your existing x86 fleet.

GB300 NVL72 puts 18 of those compute trays into one liquid-cooled cabinet, unified by NVSwitch into a single 72-GPU coherence domain. This is the rack-scale product, and that coherence domain is the only reason it costs what it costs.

HGX B300GB300 NVL72GB200 NVL72
GPUs in one coherence domain87272
Memory per GPU288 GB288 GB186 GB
Pooled memory~2.3 TB20 TB13.4 TB
Host CPUx86Grace (Arm)Grace (Arm)
Typical rack power~14 kW per chassis132–142 kW120–132 kW
CoolingAir or liquidLiquid, architecturallyLiquid, architecturally
Best fitModels that fit in 8 GPUs; enterprise inference; fine-tuningTrillion-parameter MoE, frontier training, long-context reasoningSame, at lower dense throughput

One correction worth making, since it appears in our own earlier writing and across half the internet: HGX B200 is 180 GB per GPU, not 192 GB. The 192 figure is raw package capacity. NVIDIA's own DGX B200 spec gives 1,440 GB across eight GPUs. GB200 NVL72 implies a third figure again, 186 GB, because the superchip binning differs. Three numbers for one part family, which is roughly the level of care the category currently operates at.

What actually changed between Blackwell and Blackwell Ultra

Nobody selling you a GB300 leads with this.

Same die. Same process node, TSMC 4NP. Same 208 billion transistors. Same SM count. GB200 NVL72 and GB300 NVL72 print the identical sparse NVFP4 figure: 1,440 PFLOPS per rack.

The 1.5× you keep seeing is dense-only, because dense NVFP4 went from 720 to 1,080 PFLOPS. Per GPU that works out at 15 dense against 20 sparse, a 1.33× ratio where the usual gap is 2×. FP8 didn't move at all, flat at 5 and 10 PFLOPS.

Per rackGB200 NVL72GB300 NVL72Change
NVFP4 dense720 PFLOPS1,080 PFLOPS+50%
NVFP4 sparse1,440 PFLOPS1,440 PFLOPSnone
HBM per GPU186 GB288 GB+55%
Memory bandwidth8 TB/s8 TB/snone
FP642,880 TFLOPS100 TFLOPS−97%
INT8720 POPS24 POPS−97%

Look at those last two rows. FP64 throughput collapsed by 97%, and INT8 fell by the same amount. Blackwell Ultra was never meant to be an HPC part, so if you have a computational fluid dynamics or climate modelling workload sitting in the same procurement, GB300 is the wrong answer for it.

Memory is where most of the real gain lives. A 12-Hi HBM3E stack buys 55% more of it at the same bandwidth. Special function unit throughput also doubled, 10.7 against 5 TeraExponentials per second, which matters far more than a line item on a spec sheet suggests, because softmax is exactly where long-context attention burns its clock. The higher power ceiling helps too. It lets the part actually run at its numbers.

The honest generational uplift shows up in the MLPerf v5.1 result, about 45% better offline and 25% better server throughput per GPU against GB200. Real, useful, and roughly a third of what the marketing says.

Why NVIDIA stopped selling chips and started selling racks

There is a bandwidth cliff at the edge of every server chassis, and everything about rack-scale computing follows from it.

A GPU inside an NVLink domain talks to its neighbours at 1.8 TB/s. The same GPU reaches anything outside its server through a ConnectX-7 NIC at 400 Gb/s, which SemiAnalysis put at roughly a 36× ratio in 2026. So the moment you cross the chassis boundary, your interconnect gets nearly forty times worse.

For a long time that didn't matter much, because models fit inside eight GPUs, and then they stopped fitting.

Megatron-LM's own guidance caps tensor parallelism at the server size, because tensor-parallel communication is all-to-all and murders you the moment it leaves the node. Mixture-of-experts models made this acute. The efficiency of expert-parallel inference is largely bounded by inter-device communication rather than by compute, which inverts every assumption people carry over from dense-model serving. And when your model has 256 routed experts, as DeepSeek-R1 does, you need 64 GPUs to hold them without duplication. On an eight-GPU island you fall off an 18× bandwidth cliff into RoCEv2 and the model runs badly.

SemiAnalysis measured what that costs. The same model on the same silicon runs at 4,130 tokens per second per GPU inside a large coherence domain and 941 outside it. That's a 4.39× difference produced entirely by where the wires go.

Which is why the only way to make the node bigger is to make the node the rack.

Why 72

Because that's what copper reaches.

The NVL72 backplane is passive copper, and 130 TB/s only works over sub-metre cable runs. Jensen Huang's stated rationale in 2024 was blunt. Doing it optically would have cost 20 kW of transceiver power. Copper is free and it doesn't fail, so copper reach sets the box size, box size sets the GPU count, and 72 is 18 trays of 4.

Everything else cascades from that decision. One rack means ~140 kW in one cabinet. 140 kW in one cabinet means liquid cooling, because at that density there is no other physics on offer. A rack that has to be cabled, plumbed, validated and pressure-tested as one object becomes the smallest thing that boots, and once the smallest thing that boots is a rack, the rack is what gets sold.

The actual answer to "why rack-scale" is a physical constraint on cable length, which arrived well before anyone framed it as strategy.

The counter-argument, which is real

Google reaches 9,216 TPUs in a single pod using optics and optical circuit switching, in a mechanically simpler rack. Broadcom ships scale-up Ethernet at 250 ns. NVIDIA's own documentation lists a full rack power-cycle as a recovery step, which tells you something about the blast radius when one coherent domain has a bad day.

The honest gap in the whole rack-scale case is that nobody has published measured MTBF or goodput for GB200 or GB300 NVL72 at scale. Anyone quoting you a reliability figure for these racks arrived at it by modelling, because no measurement has been published.

Most inference workloads do not need a 72-GPU domain, so if yours fits in eight GPUs, buy eight GPUs. We've made that argument at length in why Blackwell-class racks are usually overkill at the edge, and GB300 doesn't change it.

Why you can't just build one yourself

The instinct is reasonable, because seventy-two GPUs, some switches and a big rack don't read like an unsolvable engineering problem.

Every individual step is doable. The catch is that integration risk has to land somewhere, and on a self-built rack it lands on you.

A GB300 NVL72 is a plumbed pressure vessel with 1,580 kg of compute in it. Every quick-disconnect is a potential leak onto energised hardware. The NVLink backplane topology has to be validated as built, because the drawing is not the machine. Firmware across 72 GPUs, 36 CPUs, nine NVSwitch trays and eight power shelves has to be coherent, and NVIDIA revises it. The thermal loop needs hydrostatic testing and a documented flush before a single GPU sees power.

Tier-one OEMs spend enormous engineering effort producing validated systems precisely because that work is expensive and unglamorous and nobody notices it until it fails. When a coolant leak takes out a $4M rack, fault stops being the interesting question. What matters is whose warranty covers it.

The real product is the single throat to choke, and the rack is what it arrives in.

That same logic applies one level out, to the building. Coolant leaks at quick-disconnects, undersized CDUs and power delivery surprises share a common setting. They overwhelmingly happen where water, power and structure were integrated in the field, for the first time, on a live site. Which is the retrofit trap and why legacy halls fail the math.

GB300 NVL72 power requirements: the facility spec sheet

Hand this section to whoever owns your building. Every number in it is a requirement the site has to meet before the rack arrives.

Electrical

Nominal draw is 132 to 142 kW, depending on whose figure you take. NVIDIA's Enterprise Reference Architecture says "up to 142 kW." HPE's QuickSpecs give 132 kW TDP, Lenovo 135 kW. Peak power (EDPp) runs to about 155 kW, and HPE, sensibly, advises provisioning busway for 192 kW. That headroom earns its keep. AI training loads swing hard, and Uptime Institute has documented racks moving from 60 kW to 150+ kW on a one-to-two second cycle, which is why N+2 starts to look sensible where N+1 would normally do.

Inside the rack it's a single-zone 50 VDC busbar rated 1,400 A. Read that voltage twice. Most people expect 48 V, and 800 VDC only arrives with Kyber in 2027. Eight power shelves at 33 kW each, one 60 A whip at 200–277 VAC per shelf, four AC buses. Eight shelves at 33 kW is 264 kW installed against roughly 142 kW drawn, so six shelves carry the load and the arrangement reads as 6+2 shelf and 3+1 bus redundancy.

Worth knowing, and usually missed in the spec skim, GB300 carries 65 joules per GPU of electrolytic capacitance for power smoothing, which cuts peak grid draw by around 30%. Treat it strictly as smoothing. It buys you no ride-through, so your UPS sizing stays exactly where it was.

Mechanical

A populated rack is 1,580 kg sitting on 0.641 m². That works out to 2,200 to 2,470 kg/m², or 21 to 24 kN/m².

TIA-942's recommended raised floor is 1,224 kg/m². Typical raised floors, the ones you will actually be standing on, run 1,000 to 1,500. Set those against the rack and you are between 1.6 and 2.4 times over what the floor under most enterprise data halls is rated to carry. No amount of cooling redesign fixes a slab. Ten racks is 15.8 tonnes of compute, and that is before you count busbar, CDU, pumps, filters and the fluid inventory itself.

NVIDIA publishes no floor loading figure for the rack, so anyone quoting you one derived it, as we did here, and it belongs in a structural sign-off with an engineer's stamp on it.

Thermal

Lenovo splits the heat roughly 90% liquid, 10% air. At 142 kW that is about 128 kW into the loop and 14 kW still needing air handling, which surprises anyone who assumed liquid means no CRAC.

Facility water is where the good news lives. GB300 accepts supply water from 2 to 50 °C, which is ASHRAE class W45. No chiller required. In a lot of climates you will not need one at all.

Flow is strongly temperature-dependent, and it is the detail that gets missed. Lenovo publishes the curve:

Supply water temperatureFlow per rackRack pressure drop
25 °C59 L/min2.3 psi
33 °C~80 L/min
40 °C~120 L/min
45 °C177 L/min18.4 psi

Run at the 45 °C ceiling and flow triples while pressure drop goes up eightfold. Ten racks at 45 °C need 1,770 L/min. So when you size the CDU, the governing number is flow rating in N+1 mode, not kW headroom, and a unit that clears the thermal duty on paper can still fail the job.

If you're still choosing a thermal architecture rather than sizing one, we compare direct-to-chip, immersion and rear-door heat exchangers side by side.

Density and space

At 142 kW per rack, 1 MW of IT load is about seven racks. Uptime Institute's 2025 data puts the modal enterprise rack at roughly 9 kW, with more than 80% of operators running nothing above 30 kW. One GB300 rack is fifteen times the average.

The structural, electrical and thermal problems are independent of each other. Fixing one does nothing for the other two, which is where retrofit budgets go to die. Put a 48U cabinet at 1.6 tonnes on 0.64 m² and you have a slab-on-grade requirement your raised floor cannot meet, and it sits inside the full AI data center build sequence of power, cooling, network and structure.

The workloads, and what each one actually demands

"AI workload" tells you almost nothing on its own. Seven distinct things hide behind the phrase, and they want different buildings.

WorkloadBottleneckNeeds a 72-GPU domain?Latency sensitivityPower profile
PretrainingInterconnect + computeYesNoneViolently bursty
Fine-tuning (LoRA/PEFT)MemoryNoNoneSteady
Inference: prefillComputeSometimesTTFTSpiky
Inference: decodeMemory bandwidthFor large MoEPer-tokenSteady
AgenticEverything, repeatedlyUsuallyCumulativeVariable
RAGStorage + context lengthNoEnd-to-endMixed CPU/GPU
EmbeddingsMemory bandwidthNoBatch-tolerantSteady, high utilisation
Synthetic dataThroughputSometimesNoneSustained, training-like

Training has the ugliest electrical signature of the seven, and it is also the one everyone pictures. Work published by Microsoft, OpenAI and NVIDIA researchers in 2025 documented training loads oscillating across tens to hundreds of megawatts at 0.2 to 3 Hz, with GPU power floors imposed to damp it costing about 10.5% in energy overhead. SemiAnalysis puts AI data centre load volatility at roughly 15× that of conventional cloud. Grid operators have opinions about you now, and this is why.

Fine-tuning does not need this rack. LoRA and its relatives train a small adapter instead of the full weight set, which collapses memory requirements to a fraction of full fine-tuning and fits comfortably inside eight GPUs. It is also where most enterprises actually live, so if your roadmap says "customise a model on our data," this is the line you are on.

Inference splits in two, and the split determines your hardware. Prefill is compute-bound, because you process the whole prompt at once. Decode is memory-bandwidth-bound, generating one token at a time while the GPU mostly waits on HBM. Systems that disaggregate the two and run them on different hardware are winning, and McKinsey's 2026 forecast has inference growing from 20.9 to 93.3 GW at a 35% CAGR while training grows from 23.1 to 62.2 GW at 22%. Inference passes half of all AI compute before 2030.

Agentic workloads are where density earns its money. Anthropic's engineering team published that agents use about 4× more tokens than chat interactions, and multi-agent systems about 15× more — and that token usage alone explains 80% of performance variance on BrowseComp. NVIDIA instrumented a single 33-minute coding-agent session and counted 283 inference requests, 3.5 million input tokens, and context growing from 15K to 156K within the session. Every agentic user is fifteen chat users wearing a trench coat.

RAG gets talked about as a model, when what you deploy is a three-stage pipeline that embeds, retrieves and generates. Its infrastructure profile is mixed CPU and GPU, storage-heavy, and it cares about context length far more than about FLOPS. Legal discovery, clinical documentation, internal knowledge search. None of it needs a coherence domain.

Embeddings are the quiet one, bandwidth-bound and batch-friendly, small models running at enormous volume. Run an embedding fleet on GB300 and you have made a category error, an expensive one.

Synthetic data generation is technically inference and behaves like training. Throughput-bound, latency-indifferent, sustained high utilisation for days. Nemotron, Phi and Cosmos were all built this way. If you are distilling a large model into a small one, this is your workload, and the rack suits it well.

For workloads that don't need the domain, the sizing method is the one we use for translating AI servers into rack power and cooling. And the fastest-growing multi-step inference load has moved past text altogether, into vision-language-action models and the robot's local brain.

The important part: what does one rack actually deliver?

Every CEO asks the same question and almost nobody answers it. If I spend four million dollars on one rack, how many users can it serve?

Answering it means separating three quantities that usually get collapsed into a single number, and they diverge badly. Confusing them is the most expensive mistake in AI infrastructure procurement.

Capacity typeQuestion it answersWhat limits it
Memory capacityHow many sessions fit?HBM: model weights + KV cache + runtime overhead
Compute capacityHow many can be processed at once?Memory bandwidth on decode, FLOPS on prefill
Commercial capacityHow many users get an acceptable experience?Time to first token and per-token latency under load

They fall in that order, and the gap between the first and the third is roughly an order of magnitude.

Memory capacity: the arithmetic

Two things occupy HBM in any inference deployment. Model weights are a fixed cost, and KV cache grows linearly with every token in every active session.

The KV cache formula:

KV bytes per token = 2 × layers × kv_heads × head_dim × bytes_per_element

The 2 covers keys and values. Grouped-query attention collapses kv_heads, and it is the single biggest lever on this number. DeepSeek's multi-head latent attention pushes harder still, so a 671B-parameter model with MLA needs 68.6 KiB per token, less than Llama 3.1 70B's 320 KiB. The big model carries the smaller cache.

Now the case Yuri's brief asked us to check: 1M-token context sessions on Llama 4 Scout.

The first pass through the formula gave 183.1 GiB of KV cache per session. Four of those come to 732 GiB against roughly 241 GiB of usable HBM on a 288 GB GPU. Arithmetically impossible. Then we opened the config and found that Llama 4 is something other than a full-attention model. It sets no_rope_layer_interval=4, so only 12 of 48 layers do global attention across the full context. The other 36 use chunked local attention with an 8,192-token window. Rerun the arithmetic on that basis and the KV cache per session drops from 183.1 GiB to 46.9 GiB, a 3.9× difference, and the numbers work.

ConfigurationSessions per GPUSessions per rack
BF16 KV cache, weights replicated per GPU (TP=1)4.1293
FP8 KV cache, weights replicated per GPU (TP=1)8.1585
BF16 KV cache, weights sharded across the NVLink domainn/a~370
FP8 KV cache, weights sharded across the NVLink domainn/a~740

Derived figures. Assumptions: Llama 4 Scout with chunk-aware KV accounting, FP4 weights, ~84% of HBM usable after runtime overhead. Reproducible from the formula above.

Look at what happens between rows two and four. Running TP=1 keeps 72 redundant copies of the weights, which burns about 19% of the rack's HBM on duplication. Sharding across the NVLink domain gets that back. The coherence domain is the entire reason you bought this rack, and "sessions per GPU" stops being a meaningful unit the moment you use it.

Compute capacity: also about 740

The compute side surprised us. Memory capacity lands at 739 sessions, and when we ran the bandwidth-limited decode number at 20 tokens per second per user, it came out at 743.

Those two numbers land within 0.5% of each other. The rack is almost perfectly balanced, and that balance is what you get when the people who set the HBM capacity also set the bandwidth.

For measured throughput rather than modelled, MLPerf Inference v6.0 has GB300 NVL72 submissions in the Available category. One 72-GPU rack from Nebius produced 673,936 tokens per second offline and 575,580 in the server scenario on DeepSeek-R1, plus 1,046,150 / 1,096,770 on gpt-oss-120B. Worth flagging that NVIDIA's widely quoted 8,064 tokens/second/GPU comes from an eight-GPU partition, and the true rack average is 7,994.

Commercial capacity: 50 to 100

This is where the number falls off a cliff.

Prefill is the constraint here, ahead of both memory and decode. Processing a 1M-token prompt on Llama 4 Scout is 1.599 × 10¹⁷ FLOP, and 79% of that is quadratic attention cost. On the full rack that runs in 0.317 seconds, which looks fine as long as the session has the machine to itself.

Real workloads share the rack. At 100 concurrent sessions splitting it fairly, time to first token is 31.7 seconds, and at 739 sessions it stretches to 3.9 minutes. Nobody waits 3.9 minutes.

So the real answer for 1M-token context is roughly 50 to 100 concurrent sessions on a cold start, climbing toward the memory limit only with aggressive prefix caching and long-lived sessions where the expensive prefill is amortised across many turns.

Shorter contexts change everything, obviously. An 8K chatbot with 2K responses is a completely different machine, and there the rack serves thousands. Treat the specific number as incidental. What matters is that memory capacity, compute capacity and commercial capacity differ by an order of magnitude, and the vendor quoting you the first one is, strictly speaking, telling the truth.

Tokens per watt, and the honest version of the efficiency story

Cost per token is the metric everyone quotes. Tokens per watt is the one that decides whether your building can host the thing, and almost nobody publishes it.

Take the MLPerf v6.0 figure and divide: 673,936 tokens/second ÷ 142,000 W gives 4.75 tokens per second per IT watt. DeepSeek-R1 at NVFP4, offline scenario, so treat it as an upper bound. The latency-constrained server figure is 4.05.

Now push it out to the facility boundary, which is where your electricity bill lives:

BoundaryTokens/s per watt
IT load only4.75
Facility at PUE 1.104.31
Facility at PUE 1.403.39

A 21.4% swing in output per watt, decided entirely by the building. Same rack, same model, same benchmark. That is the argument for purpose-built infrastructure stated in the only unit that matters, and it is the same reasoning that opened the door to purpose-built inference silicon, which we compare in GPUs versus LPUs versus NPUs for inference.

Two efficiency claims worth retiring. The "15× cheaper tokens versus the previous generation" figure is not reproducible from any published benchmark and sits in NVIDIA's material next to an unrelated 15× ROI claim. And the "50× AI factory output" number is 10× responsiveness multiplied by 5× tokens per megawatt, which is a legitimate arithmetic operation performed on two things that shouldn't be multiplied.

The uncomfortable one: on SemiAnalysis InferenceX modelling, GB300 is about 27% more expensive per token than GB200 at 108 tokens/second/user on an 8K input, 1K output workload. GB300 wins at long context and high interactivity, and loses at short context. On sparse NVFP4 per watt it is actually a regression, 10.14 against 12.00 TFLOP/s/W, because the sparse number didn't move but the power did.

Buy it for what it's good at.

The economics, by who you are

AI cloud providers. Utilisation decides whether the rack pays back at all. Hardware-only pricing runs $3.7M to $4.0M (Loop Capital and independent analysis, 2026, and one anonymous source claims $6.0–6.5M, so treat pricing as genuinely contested). Break-even then lands between 41% and 69% on a five-year schedule. Which end you land on depends on whether you sell at year-one rates or a six-year average. Rental revenue moves as well. GB300 currently rents at $5.60 to $8.62 per GPU-hour and decays toward roughly $2.36 over six years, which makes any flat-rate payback model wrong. This is the entire engine of how the neocloud business actually works.

Depreciation is where the argument actually lives. CoreWeave's Q1 2026 shows 65% gross margin and 56% adjusted EBITDA collapsing to 1% adjusted operating margin and a $740M net loss once D&A (55% of revenue) and interest (26%) are counted. Useful life is the assumption doing the work, and nobody agrees on it. CoreWeave books six years, Amazon five, Nebius four, and Michael Burry argues for two to three. Everybody is right until the residual values print.

Colocation operators. Demand is the easy part; the hard part is whether your hall can physically accept the rack. Schneider Electric's data suggests only about one operator in five can support even 50 to 70 kW per rack, against a market average near 27 kW. For 140 kW liquid-cooled, nobody publishes a count at all. If you're contracting for GB300 space, write these into the agreement: kW per rack position, kg/m² point load, facility water supply temperature and ΔT, litres per minute per rack, and N+1 CDU capacity at flow. That choice, selling GB300 hours or leasing the kilowatts around someone else's, is a balance-sheet decision. We lay it out in GPU-as-a-service versus powered shell.

Enterprises. Most of you should not buy this rack. Fine-tuning, enterprise inference and RAG all fit inside eight GPUs. It earns its cost on trillion-parameter MoE inference, frontier training and long-context reasoning. If none of those words describe your roadmap, HGX B300 or even Hopper is the better arithmetic. Hopper is still 2026's value buy for a reason.

Sovereign programs. These buyers optimise for something else entirely. National AI initiatives buy rack-scale for reasons that have nothing to do with tokens per second and everything to do with owning the compute instead of renting intelligence. Jurisdiction is the metric that matters, and throughput is secondary.

Everyone. An allocation with no energised shell behind it is the most expensive inventory on your balance sheet. Sitting in a crate, a GB300 rack burns roughly $2,575 per day in depreciation. A 90-day slip on 1 MW costs $6.5M to $9.7M in depreciation and forgone revenue combined, between half and three-quarters of the entire shell-and-core cost of that megawatt. JLL found 57% of data centre projects slipped three months or more in 2025. That is the whole subject of what happens after the GPUs arrive.

Open-weight models: a capacity question, not a capability one

Twenty terabytes of pooled HBM changes what "we'll run it ourselves" means.

Open-weight frontier models in 2026 are large and mostly Chinese. Kimi K3 runs 2.8T parameters with 104B active, DeepSeek-V4-Pro sits at 1.6T, and GLM-5.2 ships 753B under an MIT licence. MiniMax-M3 at 428B fits on a single B300 at 4-bit, leaving about 47 GB for KV cache. And in MXFP4, Kimi K3 comes to roughly 1.49 TB, which puts the best open model in the world inside one eight-GPU HGX B300 node at 71% occupancy.

That puts the best open model in the world inside a ~14 kW chassis. Running open weights is a single-node problem. What the 142 kW rack buys you is serving those weights to a lot of people at once, with a long context.

Quantisation does not cost you much either. FP8 is effectively lossless, and NVIDIA's published NVFP4 post-training results on DeepSeek-R1 show a degradation of one point or less across seven benchmarks. The gap to closed frontier models is small and measurable, 60 against 63 on the Artificial Analysis index. Epoch AI's May 2026 reading puts open weights four months behind on capability-equivalent terms, slightly wider than in late 2025.

Whether you need a model that big is settled elsewhere. We cover what open-weight models can actually do now in detail.

What hosting GB300 actually looks like

Most vendors write this section with renders, a percentage, and the word "optimised."

Ours comes from ModulEdge's 1.5 MW liquid-cooled AI module, designed around GB300 NVL72. It holds ten GB300 racks plus eight network racks in 600 × 1,200 × 48U MGX positions, in a containerised two-storey configuration.

ParameterDesign valueBasis
IT capacity1.5 MW10 GB300 racks + 8 network racks
Design PUE1.094 annual / 1.145 design peakISO/IEC 30134-2, modelled, 12-month period
Design WUE0.037 L/kWh site, 487 m³/yearISO/IEC 30134-9, WUEsite
Coolant loopSingle loop, 43 / 53 °C, 140 m³/h, propylene glycol 25%No intermediate heat exchanger
Climate basisDubai Intl, WMO 411940, ASHRAE 2021 design conditions0.4% DB 43.3 °C, 0.4% WB 30.4 °C
Coolers selected at50 °C ambientAbove the station's recorded maximum
Spray onset41.3 °C, 127 spray hours/yearTwo coolers rather than one
White space density13.6 kW/m²1.5 MW over 110 m², from dimensioned GA drawing
Total plot1,586 m²Switchgear, transformers, 3 generators, both modules, all heat rejection
Boundary noiseLpA 59.4 dB at 10 mSelection output

White space density lands at thirteen point six kilowatts per square metre. A traditional raised-floor hall runs one to two.

The efficiency comes from one design choice that runs opposite to most builds. Instead of holding facility water at a fixed temperature, the process loop floats from 30 °C up to 43 °C as ambient rises. MGX racks accept 45 °C inlet, so the loop runs hot exactly when the outside air is hot. Approach temperature stays wide and the coolers stay dry until 42 °C outside. Dubai gets one hundred and twenty-seven spray hours in a year.

Every kelvin of approach temperature you concede has to come off the facility water, and colder facility water buys you fewer dry hours, more make-up water and more compressor energy for the rest of the asset's life. A single loop with no intermediate heat exchanger concedes nothing.

Redundancy in the cooler selection pays twice. Either cooler alone carries the full 1,559 kW duty, and running both defers spray onset from 35.6 °C to 41.3 °C, which cuts annual spray hours roughly fifteenfold. One choice delivers full-duty redundancy and a water saving.

A seawater-rejection variant reaches design PUE 1.084 with zero water, with peak and annual PUE identical, because seawater temperature barely moves through the year.

These are design and simulation figures for a named station, not measured values from an operating site, and we say so plainly. The module is designed to meet Tier III/IV principles. Uptime certifies sited facilities rather than products, so certification is something a site owner pursues after deployment. It arrives heat-reuse-ready and instrumented to report PUE, WUE, ERF and REF from day one. Article 12 of the EU Energy Efficiency Directive and Delegated Regulation 2024/1364 require exactly that of operators at 500 kW installed IT power and above.

We publish the weather station, the design conditions, the component-level power and water split, and the metric category. Fourteen vendors were reviewed while this design was being verified. Not one of them publishes all five.

What comes next, and what to build for

VR200 NVL72 lands at 190 kW in Max Q configuration and 230 kW in Max P, with a CPX variant near 370 kW. Coolant flow roughly doubles against GB300. Busbar current climbs from 1,400 A to 5,000 A and up, and the busbar has to be liquid-cooled itself. 800 VDC arrives with Kyber in 2027, not before.

So a hall built for GB300 today needs roughly 1.7 to 1.9× electrical headroom to accept VR200 later, coolant loop capacity sized for warm-water operation at 45 °C, and reserved switchgear space for an 800 VDC rectifier plant. NVIDIA's only public forward-compatibility guidance is the MGX statement that today's design investments carry forward, and that statement covers rack and system design. Your building's power and cooling sit outside it.

The full picture is in the Vera Rubin era and the 190–230 kW rack. The short version: the asset at risk isn't the chip. It's the building.

Designed for 142 kW Racks

ModulEdge builds liquid-cooled modules for GB300-class density: the floor, the busbar and the coolant loop sized for the rack you just read about, in 3 to 6 months instead of a multi-year build.

  • 5–150 kW per rack, with GB300 positions provisioned for peak draw, not nominal
  • Single-loop direct liquid cooling at 43/53 °C, ASHRAE W45 — no chiller in most climates
  • Point load, coolant flow and CDU redundancy sized on your numbers, at your supply temperature
  • Design PUE 1.094 annual, WUE 0.037 L/kWh, modelled against your site's weather station
  • Designed to meet Tier III/IV principles, instrumented for PUE, WUE, ERF and REF from day one

FAQ

How much power does a GB300 NVL72 rack use? 132 to 142 kW at nominal load, with peaks near 155 kW. HPE recommends provisioning busway capacity for 192 kW to absorb transients. Internally it runs a single-zone 50 VDC busbar rated 1,400 A, fed by eight power shelves at 33 kW each, one 60 A whip per shelf across four AC buses.

Does GB300 NVL72 require liquid cooling? Yes, the requirement is architectural. About 90% of the rack's heat goes to liquid and roughly 10% still requires air handling. There is no air-cooled configuration of NVL72. Facility water can be supplied anywhere from 2 to 50 °C, which is ASHRAE class W45 and means a chiller is often unnecessary.

How much does a GB300 NVL72 rack weigh, and what floor does it need? 1,580 kg fully populated on a 0.641 m² footprint, which is 2,200 to 2,470 kg/m² of point load. TIA-942's recommended raised floor rating is 1,224 kg/m², so most existing data halls are between 1.6 and 2.4 times under. NVIDIA publishes no floor loading figure; any number you see, including this one, is derived and needs structural sign-off.

What is the difference between GB300 and GB300 NVL72? GB300 is the superchip, Grace CPUs coupled to Blackwell Ultra GPUs on an integrated Arm platform. GB300 NVL72 is 18 of those compute trays in one liquid-cooled rack, unified by NVSwitch into a single 72-GPU coherence domain with 20 TB of pooled HBM3E. The coherence domain is what you're paying for.

Is GB300 faster than GB200? On dense NVFP4, yes, 1,080 against 720 PFLOPS per rack. On sparse NVFP4 they are identical at 1,440 PFLOPS, and FP8 is unchanged. The best-evidenced real uplift is MLPerf v5.1, roughly 45% better offline and 25% better server throughput per GPU. FP64 and INT8 throughput both fell by about 97%, so GB300 is not an HPC part.

How many users can one GB300 NVL72 rack serve? It depends entirely on context length, and on which capacity you mean. At 1M-token context on Llama 4 Scout, roughly 740 sessions fit in memory with FP8 KV cache and sharded weights, but prefill cost limits acceptable-latency service to about 50 to 100 concurrent sessions on a cold start. At short context the rack serves thousands. Memory capacity, compute capacity and commercial capacity differ by roughly an order of magnitude.

How many tokens per second does a GB300 NVL72 rack produce? In MLPerf Inference v6.0, one 72-GPU rack produced 673,936 tokens/second offline and 575,580 in the server scenario on DeepSeek-R1, and 1,046,150 / 1,096,770 on gpt-oss-120B. That works out to about 4.75 tokens per second per IT watt, falling to 4.31 at facility PUE 1.10 and 3.39 at PUE 1.40.

How many GB300 racks fit in 1 MW? About seven at 142 kW nominal. A 1.5 MW module supports ten GB300 racks plus supporting network racks, at roughly 13.6 kW per square metre of white space compared with 1 to 2 kW/m² for a traditional raised-floor hall.

Can a colocation facility host GB300 NVL72? Most facilities cannot host it today. Schneider Electric's data suggests only about one operator in five supports even 50 to 70 kW per rack, against a market average near 27 kW. Before signing, specify kW per rack position, kg/m² point load, facility water supply temperature and ΔT, litres per minute per rack at your design temperature, and N+1 CDU capacity rated on flow instead of kW.

Should my company buy a GB300 NVL72? Probably not, for most companies. The rack earns its cost on trillion-parameter mixture-of-experts inference, frontier training, and long-context reasoning workloads. Fine-tuning, RAG, embeddings and standard enterprise inference all fit comfortably in eight GPUs. HGX B300 or previous-generation Hopper silicon is usually the better arithmetic.

Yuri Milyutin

Managing Partner at ModulEdge