Key takeaways

  • This is not a rack-for-rack comparison. NVIDIA GB300 NVL72 is a 72-GPU rack-scale system. Google documents Ironwood as TPU7x slices that can scale to a 9,216-chip pod.
  • Per accelerator, TPU7x has 192 GiB of HBM and Blackwell Ultra has 288 GB of HBM3E. NVIDIA lists 20 TB of GPU memory for the full GB300 NVL72 rack.
  • The fabrics use different boundaries and units. Google lists 1,200 GB/s of bidirectional ICI bandwidth per TPU7x chip; NVIDIA lists 1.8 TB/s per GPU and 130 TB/s aggregate NVLink bandwidth across a GB300 NVL72 rack.
  • Google's 1.77 PB figure is an aggregate superpod number. It should not be read as a single, flat 1.77 PB memory address space available to every operation.
  • Both are deployable products, not paper launches. Google documents TPU7x use through Compute Engine and GKE, while NVIDIA publishes GB300 system documentation and points customers to its sales channel.

Why the comparison needs careful boundaries

Google Ironwood and NVIDIA GB300 NVL72 are both designed for large AI workloads, but they describe different layers of infrastructure. Ironwood is an accelerator generation offered through Google Cloud. GB300 NVL72 is a complete liquid-cooled rack built around 72 Blackwell Ultra GPUs and 36 Grace CPUs.

That difference matters. Comparing a 9,216-chip Ironwood pod with one 72-GPU GB300 rack can illuminate architectural choices, but it cannot establish which platform is faster, cheaper or more efficient. Those conclusions require the same model, software, precision, batch size and system boundary.

Official specifications at a glance

MetricGoogle Ironwood (TPU7x)NVIDIA GB300 NVL72
Accelerator memory192 GiB HBM per chip288 GB HBM3E per GPU
HBM bandwidth7,380 GB/s per chipNVIDIA lists 20 TB GPU memory and up to 576 TB/s aggregate bandwidth per rack
Scale-up systemUp to 9,216 chips per pod72 GPUs and 36 Grace CPUs per rack
Scale-up fabric1,200 GB/s bidirectional ICI per chipFifth-generation NVLink, 1.8 TB/s per GPU and 130 TB/s aggregate per rack
Scale-out network100 Gbps DCN bandwidth per chipConnectX-8 provides 800 Gb/s per GPU
Peak low-precision compute4,614 FP8 TFLOPS per chip720 PFLOPS FP8/FP6 per rack, with sparsity

The units above follow each vendor's documentation. They are useful for sizing memory and network requirements, but they are not normalized benchmark results.

Memory: local capacity versus aggregate capacity

Google's TPU7x documentation lists 192 GiB of HBM and approximately 7.38 TB/s of HBM bandwidth per chip. The same documentation explains that an Ironwood chip contains two chiplets, each with its own 96 GB memory space; communication between chiplets is handled with collective operations.

At pod scale, Google describes 1.77 PB of shared HBM across 9,216 chips. That is an aggregate system figure. It does not erase the software work involved in partitioning model weights, activations and communication across thousands of devices.

NVIDIA lists 288 GB of HBM3E on each Blackwell Ultra GPU and 20 TB of GPU memory across GB300 NVL72. NVIDIA also lists 37 TB of “fast memory” for the rack, a broader number that includes GPU and CPU memory. Keeping those two figures separate avoids turning system memory into GPU HBM.

For buyers, the practical question is not simply which total is larger. It is whether a workload fits within one accelerator, one tightly connected scale-up domain, or a larger cluster that adds scale-out traffic.

Google lists TPU7x bidirectional Inter-Chip Interconnect bandwidth at 1,200 GB/s per chip, equivalent to 9.6 Tb/s. A 9,216-chip pod combines ICI with optical circuit switching. Google's current reliability documentation says the pod is organized as 144 cubes of 64 chips, with optical switching connecting those cubes.

GB300 NVL72 uses fifth-generation NVLink inside the rack. NVIDIA documents 1.8 TB/s per GPU and 130 TB/s of aggregate, non-blocking all-to-all bandwidth for the 72-GPU NVLink domain. Scale-out traffic leaves that rack through ConnectX-8 networking at 800 Gb/s per GPU using InfiniBand or Spectrum-X Ethernet.

Those bandwidth figures should not be ranked as if they describe the same link. Google's number is per TPU chip, while NVIDIA publishes both per-GPU and rack-aggregate NVLink figures. The physical topology and the software collectives determine how much of that bandwidth an application can use.

Deployment and software choices

TPU7x is available through Google Compute Engine and Google Kubernetes Engine. Google supports JAX and PyTorch on TPU7x, while its current documentation says TensorFlow is not supported. Teams choosing Ironwood are therefore choosing Google Cloud's TPU hardware, XLA-oriented software path and reservation model together.

GB300 NVL72 is a rack-scale platform sold through NVIDIA and system partners. Its standard environment centers on CUDA and NVIDIA's networking and management stack. NVIDIA's product documentation positions the system for both training and reasoning inference, and its published reference architecture describes the rack components and scale-out network in detail.

Neither vendor's peak specification answers an application-level cost question. A credible evaluation needs delivered tokens per second or training throughput, model quality, utilization, power at the chosen system boundary and the actual cloud or procurement price.

What infrastructure planners can conclude

Ironwood offers a much larger documented scale-up domain: up to 9,216 TPU chips under Google's pod architecture. GB300 NVL72 packages a smaller 72-GPU scale-up domain into a defined rack with standardized NVIDIA networking for building larger clusters.

Memory totals and fabric bandwidth indicate the kinds of models each system is built to support. They do not establish a universal winner. The defensible shortlist comes from workload portability, framework support, capacity access and measured performance on the buyer's own model—not from dividing one vendor's pod total by another vendor's rack total.

Sources

Sources

This article was researched and fact-checked against the following sources: