August 30, 2026
OpenAI's Jalapeño Chip: What 160 kW in Two Racks Means for Your Floor Plan
OpenAI's Jalapeño beat GB300 on tokens per watt. SemiAnalysis, who ran the test, call that the wrong comparison. The right one, and the 160 kW nobody wrote down.

On 25 August 2026 SemiAnalysis published measured results for Jalapeño, the inference accelerator OpenAI designed with Broadcom. The chart that travelled compares it against GB300. On that chart, a 700-watt part beats a 1,400-watt one.
SemiAnalysis ran the benchmark, and they say that comparison does not hold up, because the two chips use different memory generations. Jalapeño has HBM4. Blackwell has HBM3E. The like-for-like part is Vera Rubin, which also uses HBM4.
Jalapeño still wins that one. The number that should interest anyone building the room around it is neither 700 nor 1,400. It is 160 kW.
What was actually measured, and against what
Jalapeño is an inference chip. Inference means running a trained model rather than training one. It was tested on InferenceX, a public benchmark from SemiAnalysis, across GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. OpenAI published its own account the same day.
Read the denominator before the ratios, because it is unusual and it matters. The headline metric is output token throughput per all-in utility megawatt. Not per chip, per watt at the meter. As SemiAnalysis point out, tokens per second per megawatt reduces to tokens per joule. It measures how well a system turns energy into product.
Three caveats travel with all of it.
The numbers were supplied by OpenAI. SemiAnalysis verified runs in person in the lab, but did not run the full suite. There are no AgentX results yet, and that is the harder multi-turn, long-context test. The workload used is a straightforward one to tune for, and the models are not the largest open ones available.
Jalapeño posted those results using single token prediction, with no speculative decoding and no prefill-decode disaggregation. Rubin's comparison figures used multi-token prediction. So the winning side was working with optimizations switched off that the losing side had switched on.
Availability runs the other way, though. Vera Rubin systems are shipping to customers now. Jalapeño is at engineering-sample stage, with production scheduled to ramp across 2027. On cost per token, SemiAnalysis put the two level.
Power is the constraint, and now both sides say so out loud
An operator selling inference converts electricity into product. We have argued for years that the metric deciding hardware is output per watt, not cost per chip. What we have never had is a number to put behind it.
There is one now, and the people saying it loudest are not us.
At Computex 2026, NVIDIA's chief executive put it plainly: "If you have 1 gigawatt of power, then throughput per watt is revenue." NVIDIA said the same at Hot Chips 2026. The data center is power limited today. SemiAnalysis make the same argument from the operator's side. Nobody can simply acquire more megawatts, because adding accelerators and adding grid capacity happen on entirely different timescales. So operators build generation behind their own meter rather than wait for an interconnection. We have written up what that looks like in practice separately.
So the metric is now common ground. What we still cannot tell you is which silicon wins on it. We hold no published inference-efficiency figure of our own. Nor do we publish one. Anyone in our category asserting a tokens-per-watt superlative for their preferred accelerator is inventing it.
One precision worth holding on to, because it gets confused constantly. PUE is total facility energy divided by IT energy, and it measures the overhead a facility adds around the compute. About how much useful work that compute produced it says nothing at all. ISO/IEC 30134-2 says so directly. Do not use it to compare data centers.
Two facilities can post identical PUE and differ enormously in the tokens they deliver. That difference lives in the silicon and the software.
160 kW, in two racks, in a sixteen-rack block
Here is the part almost nobody covering this story wrote down.
A Jalapeño system is not a rack. There are two: a host rack of sixteen CPU trays, and beside it an accelerator rack. That one holds sixteen trays of eight ASICs each, so 128 accelerators, plus eight switch trays. System-level design is with Celestica.
Then the topology. All 128 accelerators connect over a copper backplane inside the rack. Above that sits a global domain of 2,048 accelerators spanning sixteen racks, stitched together with a mix of copper and optics. Passive copper alone runs to 8,192 differential pairs per rack.
For anyone designing the space, that changes three things at once.
The deployable unit is a pair of racks, not a rack, so floor pitch and power distribution have to be laid out in pairs. The scale-up domain is a contiguous sixteen-rack block. That is a floor-plan constraint, not a chip constraint, and it will not tolerate being split across a hall for convenience. Then the host rack, provisioned at roughly 50 kW and drawing 31 kW. Read that as the headroom argument in miniature. Someone chose to pay for capacity they are not using, deliberately, because the alternative is a rebuild.
All three get decided in the drawing, months before any hardware arrives.
The forward-compatibility question is the one a technical buyer should be asking, and part of it has a published answer. MGX is NVIDIA's 600 mm rack architecture. NVIDIA has stated that the same MGX footprint supporting GB300 NVL72 will support Vera Rubin NVL72, Vera Rubin NVL72 CPX and Vera Rubin CPX. A module built around that footprint does not strand when the next NVIDIA generation lands.
What Jalapeño adds is that MGX is no longer the only geometry that matters. A two-rack system with its own switch trays, its own backplane and a sixteen-rack scale-up domain is a different physical problem. A different integrator is building it. Two caveats we would rather say than have you find. Footprint continuity is a different thing from power continuity, and a rack that physically fits may still be a rack your electrical and thermal design cannot feed. Also: NVIDIA publishes no rack power figure for Vera Rubin NVL72. Every number in circulation for it comes from analyst supply-chain reporting rather than a manufacturer specification. Treat it accordingly.
Efficiency does not move the compliance line
This is the part that surprises European buyers, and a more efficient chip makes it more relevant rather than less.
The EU reporting obligation starts at 500 kW of installed IT power demand, under Commission Delegated Regulation (EU) 2024/1364, made under the Energy Efficiency Directive (EU) 2023/1791. Reports are due annually on 15 May. Set that against a single Jalapeño accelerator rack at 130 kW.
Four racks. Most buyers we meet file "reporting threshold" under somebody else's problem. Here is why that instinct fails. The threshold is written on installed IT power, not on work delivered. A more efficient accelerator raises the output you get from those watts, and the installed power stays where it was. Efficiency changes what you get for being in scope. It will not take you out of it.
Above roughly 1 MW of total rated energy input a second obligation lands. Waste heat must be utilized unless that is shown not to be technically or economically feasible, and new or substantially refurbished facilities must carry out a cost-benefit analysis. The instrument is EED Article 26(6) and 26(7).
The EU itself sets no fixed percentage, and that trips people up constantly. The percentages that bind are national. Germany's Energieeffizienzgesetz § 11 requires an energy reuse factor of at least 10% from 1 July 2026, rising to 20% by 2028. Those numbers are moving. A Novelle cleared cabinet on 24 June 2026 proposing a higher entry threshold and relaxed ceilings. As at 29 August 2026 it has not entered force. Re-check before you rely on them.
None of this retrofits cheaply. Return-water temperature, hydraulic interface, space for a heat exchanger, sub-metering fine enough for PUE2 or PUE3 reporting: all of it is decided before the module is built. Or it is never decided. What a heat sink actually has to accept, and at what temperature, is covered in our piece on data center heat reuse.
We describe what the instruments say and what they mean for how a module is designed. We do not give legal or permitting advice, and compliance attaches to an operator and a site, never to a shipped product.
The question this leaves open
Read the last section of the SemiAnalysis piece and the story stops being about silicon. OpenAI's stated next target for Jalapeño is 100 MW, it is partnering with neoclouds to deploy, and it is gathering reliability data with data center partners. SemiAnalysis judge that the remaining hurdles are mostly hardware. How much can be produced. Whether it deploys, operates, gets monitored properly and holds up over time. Software, in their reading, is the proven part.
That is our whole argument, made by someone else, about someone else's chip. We did not need it to be true and it is inconvenient to have it arrive from outside, but there it is.
The open question is the one underneath. If inference silicon keeps getting more efficient per watt, and the workload keeps shifting toward many smaller deployments rather than a few enormous training runs, does the case for gigawatt-scale single campuses still hold?
We do not know. Nobody has answered it for us, and we would rather publish the question than pretend to a position we cannot defend. It is also the most interesting argument available in this category. Almost everyone in it is certain instead. The strongest version of the small-and-many case we have seen so far is the national one, which we set out in sovereign AI infrastructure.
Two clocks are worth holding side by side while you think about it. They measure entirely different things and they are not a ratio. Treat the pairing as a sense of scale. The IEA reported in November 2025 that grid connection queues run two to ten years across the EU. In the FLAP-D hubs the average is seven to ten. OpenAI taped out Jalapeño in November 2025 and had it running benchmarks nine months later.
The slow clock decides where compute can physically exist. Nobody writes about that one, though we have made a start on what operators do instead of waiting, in on-site power.
What to ask before you fix a design
Six questions, in the order they matter:
- What is the deployable unit for the hardware you expect to install: a rack, or a rack pair with a separate host?
- How large is the scale-up domain, and does the floor plan hold that many racks contiguously?
- What power per rack position does the electrical topology actually supply, as distinct from what today's hardware draws?
- What is the return-water temperature, and what would a heat-recovery tie-in physically require?
- Is the module sub-metered for what the operator will be legally required to report, at the threshold that applies to them?
- What specifically happens to this module if the accelerator inside it changes generation?
Questions three and four get answered at the site, not in a brochure. Our note on site selection and preparation covers what a survey has to establish first.
You can change the silicon later. The power path, the water temperatures, the rack pitch and the metering stay where you put them. Fix those first, and let the chip argument happen above them.
