On this page · 13 sections
- The two spec sheets, as each vendor publishes them
- Trap 1: "GPU" does not mean the same thing in both quotes
- Trap 2: FP4 and NVFP4 are not the same number
- Trap 3: rack totals are marketing arithmetic, not measured throughput
- Trap 4: one of these racks is not a product
- Trap 5: the scale-up bandwidth numbers collide at exactly the same figure
- Trap 6: memory capacity is the one clean comparison
- Trap 7: the software stack is where the real switching cost sits
- What to put in the RFP
- India-specific considerations
- FAQ
- How eCorpIT can help
- References
Summary. AMD's Helios page lists 72 Instinct MI455X GPUs, 2.9 exaFLOPS of FP4, 1.4 exaFLOPS of FP8, 31 TB of HBM4 and 260 TB/s of scale-up bandwidth per rack. NVIDIA's Vera Rubin deep dive, published 5 January 2026, lists 72 Rubin GPUs per NVL72 system, 288 GB of HBM4 and 22 TB/s per GPU, 50 PFLOPS of NVFP4 inference per GPU, and 3.6 TB/s of NVLink 6 per GPU. Both platforms reach volume in the second half of 2026. Almost none of those numbers can be placed side by side without being restated first: the accelerator counts use different units, the FP4 figures use different datatypes and different measurement conditions, and one of the two racks is not a product you can buy. For scale, AWS raised EC2 Capacity Block rates on 1 July 2026 to $14.04 per accelerator-hour for P6-B300, which puts a 72-accelerator rack-equivalent at roughly $24,261 a day at list. Here are the seven places the two spec sheets stop being comparable, and what to ask instead.
The two spec sheets, as each vendor publishes them
Both numbers below are quoted from the vendors' own pages, not from press coverage. That matters, because the secondary write-ups of both platforms garble the units badly.
AMD publishes rack-level figures on its Helios product page, every one of them carrying footnote 1: "Based on AMD internal analysis, actual results subject to change. MI450-001." NVIDIA publishes per-GPU and per-tray figures in a technical blog by Kyle Aubrey, and publishes no rack-level exaFLOPS number at all in that document.
| Metric | AMD Helios (vendor-published) | NVIDIA Vera Rubin (vendor-published) |
|---|---|---|
| Accelerators per rack | 72x Instinct MI455X | 72 Rubin GPUs per NVL72 |
| HBM4 per accelerator | Up to 432 GB | Up to 288 GB |
| Memory bandwidth per accelerator | 19.6 TB/s | Up to 22 TB/s |
| Low-precision compute | 2.9 EF FP4 per rack; up to 40 PFLOPS FP4 per GPU | 50 PFLOPS NVFP4 inference per GPU; 35 PFLOPS NVFP4 training per GPU |
| HBM4 per rack | 31 TB | Not stated |
| Scale-up fabric | 260 TB/s aggregate, UALink over Ethernet | NVLink 6, 3.6 TB/s per GPU |
| Scale-out fabric | 43 TB/s, Pensando Vulcano NICs | ConnectX-9 at 1,600 Gb/s per GPU |
| Host CPU | EPYC "Venice", up to 256 cores, 1.6 TB/s | Vera, 88 Olympus cores, 176 threads, 1.2 TB/s |
| Rack power draw | Not stated | Not stated |
| Availability | Volume deployments expected 2H 2026 | Not stated in the platform deep dive |
Two things jump out of that table before any analysis. Neither vendor publishes a rack power figure on these pages, and the compute rows are written in different units by design.
Trap 1: "GPU" does not mean the same thing in both quotes
AMD counts 72 MI455X GPUs per Helios rack. NVIDIA counts 72 Rubin GPUs per NVL72. So far the units agree.
The confusion arrives from the naming. NVIDIA's own comparison table gives the Rubin GPU 336 billion transistors across 2 compute dies, the same die count as Blackwell's 208 billion transistors across 2 dies. Multiply 72 GPUs by 2 dies and you get 144, which is where the "NVL144" label circulating in coverage comes from. The deep dive itself uses NVL72 throughout and describes a superchip as "two Rubin GPUs with one Vera CPU," with two superchips per tray.
If a quote in front of you says NVL144, that is not a different, larger rack. Ask the vendor to restate the configuration as packages, as compute dies, and as scheduler-visible devices, because those three numbers diverge and only the last one determines how many workers your orchestrator will see. Teams that have already been through this with Kubernetes device plugins will recognise the problem from the DRA GA migration for GPU ResourceClaims.
Trap 2: FP4 and NVFP4 are not the same number
AMD's Helios page says the platform supports "both OCP and MX datatypes" and quotes 2.9 EF of FP4 and 1.4 EF of FP8 per rack. NVIDIA quotes NVFP4, its own four-bit format, at 50 PFLOPS per GPU for inference and 35 PFLOPS per GPU for training.
Read the footnotes on NVIDIA's table carefully. The 50 PFLOPS inference figure is marked "Transformer Engine compute." The 35 PFLOPS training figure is marked "Dense compute." Only the training number is explicitly labelled dense. The inference number is qualified by a hardware block, not by a sparsity condition, and the document never states whether that 50 PFLOPS is dense or sparse.
That single missing word is the difference between a 2x error and a correct model. Any procurement spreadsheet that divides a headline FP4 number by a price is wrong unless the sparsity assumption is written down next to it. The same discipline applies further down the stack, as the arithmetic in our H100 FP8 versus BF16 cost per million tokens analysis showed: the datatype choice moves cost per token more than the hardware generation does.
Trap 3: rack totals are marketing arithmetic, not measured throughput
AMD's rack figure of 2.9 EF FP4 is 72 multiplied by the per-GPU 40 PFLOPS, rounded. Its own FAQ states both numbers in the same sentence, so the derivation is transparent.
NVIDIA does not publish a rack FP4 total in the platform deep dive. Doing the same multiplication yourself gives 72 by 50 PFLOPS, or 3.6 EF of NVFP4 inference, and 72 by 35 PFLOPS, or 2.52 EF of NVFP4 training. Those are derived figures, not NVIDIA claims, and they inherit every ambiguity from trap 2.
Neither total is a throughput measurement. Both are peak silicon arithmetic with no allowance for memory stalls, collective communication, or the scheduler. The gap between peak and sustained on real inference serving is large enough that we would treat any rack-level exaFLOPS number as a bound, not a forecast. Model the workload instead: our B200 versus H100 inference cost per token work found the useful comparison sits at tokens per second per dollar, never at peak FLOPS.
Trap 4: one of these racks is not a product
This is the trap most likely to survive into a signed contract, and AMD states it plainly in its own FAQ: "It's a reference design, not a product for sale. The AMD Helios Rackscale solution is a blueprint for OEM and ODM partners to build their own branded systems based on open ORW standards."
So a Helios quote does not come from AMD. It comes from HPE, Celestica, or another integrator building to the Open Rack Wide specification that Meta submitted to the Open Compute Project. AMD's published 2.9 EF and 31 TB describe the reference design. What lands on your floor is a partner's interpretation of it, and the integration, firmware, and support terms are that partner's, not AMD's.
That is not a weakness of the open model. It is a different contracting shape, and it changes who you escalate to at 3am. Get the performance numbers restated in the OEM's own paper, under the OEM's own footnote, before those figures reach a business case.
| What the number describes | AMD Helios | NVIDIA Vera Rubin |
|---|---|---|
| Who publishes it | AMD, as a reference design | NVIDIA, as a platform |
| Who sells you the rack | OEM or ODM partner | NVIDIA and its system partners |
| Rack standard | OCP Open Rack Wide, double-wide | NVIDIA rack design |
| Scale-up fabric standard | UALink over Ethernet, UEC | NVLink 6, proprietary |
| Software stack | ROCm, open source | CUDA and NVIDIA libraries |
| Footnote on headline specs | "Based on AMD internal analysis" | Per-GPU figures, table footnotes on datatype |
Trap 5: the scale-up bandwidth numbers collide at exactly the same figure
AMD calls 260 TB/s of aggregate scale-up bandwidth "industry leading" on the Helios page. NVIDIA's platform deep dive carries the figure "260 TB/s per NVL72 rack" in an image caption.
Those are the same number for the same unit of infrastructure. One of two things is true: the two fabrics genuinely land in the same place, or the two vendors are counting directionality differently and the figures only look identical. NVIDIA's interconnect table lists NVLink at 3,600 GB/s per GPU explicitly marked bi-directional; AMD's page does not state directionality for its 260 TB/s.
Ask for it in writing. A bidirectional aggregate compared against a unidirectional one is a 2x error hiding inside a number that reads as a tie.
The scale-out comparison is worse, because the two vendors do not even use the same base unit. AMD publishes 43 TB/s of rack scale-out and claims "over 50% more than competition." NVIDIA publishes 1,600 Gb/s per GPU through ConnectX-9 and no rack aggregate. Bits against bytes, per-GPU against per-rack: there is no honest way to check AMD's competitive claim from NVIDIA's own primary document, so treat it as unverified until both vendors publish the same unit.
Trap 6: memory capacity is the one clean comparison
After five traps, one row of the table survives contact with scrutiny.
AMD publishes up to 432 GB of HBM4 per MI455X and 31 TB per rack. NVIDIA publishes up to 288 GB of HBM4 per Rubin GPU. Both are capacities, both are stated in the same unit, and capacity does not depend on a measurement condition. On a like-for-like 72-accelerator rack, AMD's own figure of 31 TB compares against a derived 20.7 TB for NVIDIA.
That 50 percent capacity advantage is the most defensible claim either vendor makes, and it maps directly onto a workload property you can test: how many parameters and how much KV cache fit before you are forced into tensor parallelism across the fabric. For long-context inference that is frequently the binding constraint, as the trade-offs in our 2M-token long context versus RAG comparison show.
Bandwidth goes the other way. NVIDIA publishes up to 22 TB/s per GPU against AMD's 19.6 TB/s, about 12 percent higher per accelerator.
Trap 7: the software stack is where the real switching cost sits
AMD's page lists ROCm with support for PyTorch, TensorFlow, JAX, ONNX Runtime, vLLM, and Triton, and describes the platform as avoiding vendor lock-in. NVIDIA's platform includes six chips, with the Vera CPU, Rubin GPU, NVLink 6 switch, ConnectX-9, BlueField-4 DPU, and Spectrum-6 Ethernet switch, and a seventh added on 16 March 2026.
Framework support is table stakes now. The cost is not in whether PyTorch runs. It is in the kernels your team has already written, the inference server you have tuned, the profiler your engineers know, and the six months of operational muscle memory that does not transfer. The real cost of a platform move is usually the migration, not the code.
Price that explicitly. A capacity decision made on a 10 percent FLOPS difference will be reversed by a two-quarter software ramp, and the same pattern appears every time GPU spend is examined seriously, as our review of GPU spend as the top FinOps concern in 2026 found.
What to put in the RFP
Seven questions, each aimed at one trap above. All seven should be answerable in writing by any serious vendor.
- State the accelerator count as packages, as compute dies, and as scheduler-visible devices.
- For every FLOPS figure, state the datatype, whether it is dense or sparse, and which hardware block produced it.
- Provide sustained tokens per second on a named model at a named batch size and context length, not peak FLOPS.
- Name the legal entity selling the rack and the entity carrying the performance warranty.
- State whether every bandwidth figure is unidirectional or bidirectional, and in which unit.
- State rack power draw at sustained load, in kW, and the inlet water temperature the design assumes.
- State the delivery quarter and the penalty for slipping it.
Point six deserves emphasis because neither vendor answers it on the pages examined here. AMD describes a power shelf, a vertical busbar, and a cooling manifold with quick-disconnect couplings, and refers to "power-constrained datacenters," but publishes no kW figure. NVIDIA's deep dive contains no kilowatt figure either. For anyone signing a colocation contract, that is the number that decides whether the rack can be installed at all.
India-specific considerations
The India path to this hardware is already named. AMD and Tata Consultancy Services announced on 16 February 2026 that TCS, through its subsidiary HyperVault AI Data Center Limited, will co-develop a Helios-based rack-scale design supporting up to 200 MW of capacity in support of India's national AI initiatives.
Dr. Lisa Su, Chair and CEO, AMD, said: "AI adoption is accelerating from pilots to large-scale deployments, and that shift requires a new blueprint for compute infrastructure. With 'Helios,' we are delivering an open, rack-scale AI platform designed for performance, efficiency, and long-term flexibility. Together with TCS, we are enabling enterprises across India to deploy AI at scale today while building the compute foundation of tomorrow."
K. Krithivasan, MD and CEO, TCS, said: "This collaboration lays the foundation for AMD's first 'Helios' powered AI infrastructure in India."
For an Indian enterprise, three consequences follow. First, capacity of this class will be available domestically rather than only through a hyperscaler region, which changes the data residency conversation under the Digital Personal Data Protection Act 2023, a point developed further in our work on India's sovereign AI and the IndiaAI Mission. Second, neither AMD nor TCS has published a rupee capex figure or a per-GPU-hour rate for the HyperVault build, so any business case using an Indian price today is using an estimate, and should say so. Third, state-level incentives materially change the landed cost, and the terms differ by state, as the Karnataka data centre policy for GCC and AI infrastructure shows.
Most Indian teams reading this will not buy a rack in 2027. They will rent capacity, which makes the relevant comparison the hourly rate and the egress bill rather than the spec sheet. AWS raised EC2 Capacity Block prices for ML GPU instances by roughly 20 percent effective 1 July 2026, to $14.04 per accelerator-hour for P6-B300 and $12.355 for P6-B200, with P5 in US regions at $5.191. Those are the numbers that will actually appear on an Indian AI budget in 2027, and the data movement around them is priced separately, as our cloud egress fee comparison for AI inference sets out.
FAQ
How eCorpIT can help
eCorpIT is a CMMI Level 5 certified technology organisation in Gurugram, and our senior engineering teams work with AWS, Microsoft and Google platforms on AI infrastructure and cost engineering. We help enterprise teams normalise vendor spec sheets into a comparable model, build the sustained-throughput benchmarks that peak FLOPS numbers cannot substitute for, and structure the RFP questions above so the answers arrive in writing. If you are sizing 2027 AI capacity and want the arithmetic checked before it reaches a board paper, contact us.
References
Last updated: 22 July 2026.