Key takeaways
- NVIDIA HGX B300 is an eight-GPU platform for OEM server systems. GB300 NVL72 is a complete rack-scale system with 72 Blackwell Ultra GPUs and 36 Grace CPUs.
- Both use B300-class Blackwell Ultra GPUs with 288 GB of HBM3e per GPU, but the systems expose very different scale-up domains: eight GPUs for HGX B300 and 72 GPUs for GB300 NVL72.
- HGX B300 provides 2.30 TB of GPU memory and up to 64 TB/s of aggregate GPU-memory bandwidth per node. GB300 NVL72 provides 20 TB and up to 576 TB/s across the rack.
- NVIDIA lists 14.4 TB/s of total aggregate NVLink bandwidth for an HGX B300 baseboard and 130 TB/s for the GB300 NVL72 rack.
- GB300 NVL72 is liquid-cooled and NVIDIA’s current reference architecture says the full rack can require up to 142 kW. HGX B300 power and cooling depend on the server built around the baseboard.
The short answer
HGX B300 and GB300 NVL72 use the same Blackwell Ultra GPU generation, but they are not two sizes of the same server. HGX B300 is an eight-GPU baseboard that an OEM integrates with host CPUs, memory, storage, networking, power and cooling. GB300 NVL72 is a rack-scale Grace Blackwell system whose 72 GPUs share one in-rack NVLink domain.
That distinction is more useful than a single peak-performance number. It determines how much accelerator memory is available inside one scale-up domain, which CPUs are attached, how the system connects beyond that domain and what the facility must support.
| Specification | NVIDIA HGX B300 | NVIDIA GB300 NVL72 |
|---|---|---|
| Blackwell Ultra GPUs | 8 | 72 |
| GPU memory | 2.30 TB HBM3e | 20 TB HBM3e |
| Memory per GPU | 288 GB HBM3e | 288 GB HBM3e |
| Aggregate GPU-memory bandwidth | Up to 64 TB/s | Up to 576 TB/s |
| NVLink domain | 8 GPUs on one baseboard | 72 GPUs across one rack |
| Aggregate NVLink bandwidth | 14.4 TB/s | 130 TB/s |
| CPU design | Chosen by the OEM system builder | 36 NVIDIA Grace CPUs |
| Cooling and power | System-dependent | Liquid-cooled; up to 142 kW for the full rack |
The figures above come from NVIDIA’s current HGX B300 Enterprise Reference Architecture, GB300 NVL72 product specifications and NVL72 reference architecture.
Start with the B300 GPU
NVIDIA lists 288 GB of HBM3e per B300 GPU and up to 8 TB/s of memory bandwidth per GPU. Those are the consistent building-block figures behind both systems.
The aggregate memory figures follow the system counts. Eight GPUs produce 2,304 GB, which NVIDIA rounds to 2.30 TB for HGX B300. Seventy-two GPUs produce 20,736 GB, which the GB300 product page reports as 20 TB. Likewise, the published aggregate memory-bandwidth figures are up to 64 TB/s for the eight-GPU HGX node and up to 576 TB/s for the 72-GPU rack.
This also resolves a common specification mismatch. Some third-party pages quote 279 GB for B300, but NVIDIA’s current product and architecture pages specify 288 GB. For procurement or capacity planning, the manufacturer’s current platform documentation is the safer reference.
What HGX B300 includes
The HGX B300 baseboard connects eight Blackwell Ultra GPUs with fifth-generation NVLink and NVSwitch. NVIDIA specifies 14.4 TB/s total aggregate NVLink bandwidth and 1,800 GB/s GPU-to-GPU bandwidth for the baseboard.
HGX is the accelerated-computing foundation rather than a complete fixed rack. An OEM builds a server around it, adding host CPUs, system memory, local storage, BlueField-3 and ConnectX-8 networking, power supplies and cooling. NVIDIA’s reference design calls for eight ConnectX-8 SuperNICs, preserving a one-to-one GPU-to-NIC relationship.
That flexibility matters. Buyers can select a certified server and then scale out multiple HGX nodes over InfiniBand or Ethernet. The trade-off is that communication outside each eight-GPU NVLink domain travels over the external compute network.
What GB300 NVL72 adds
GB300 NVL72 integrates 72 Blackwell Ultra GPUs and 36 Grace CPUs in one liquid-cooled rack. NVIDIA’s reference design divides the rack into compute trays, each with four GPUs and two Grace CPUs. Eighteen such trays supply the rack totals.
Nine NVLink Switch trays connect the GPUs into one 72-GPU domain with 130 TB/s of aggregate NVLink bandwidth. The rack also includes ConnectX-8 adapters for scale-out compute traffic and BlueField-3 DPUs for the north-south service network. Our separate GB300 NVL72 networking guide maps those fabrics and their boundaries.
The rack is also a facility-level commitment. NVIDIA’s reference architecture specifies liquid cooling, eight 33 kW power shelves and up to 142 kW for the full rack. That is not the same as saying the GPUs continuously consume 142 kW: it is the published full-rack infrastructure requirement, including the integrated system.
How to interpret the memory advantage
The 72-GPU rack offers nine times the GPU count and aggregate HBM capacity of one HGX B300 node. It also keeps those GPUs inside one NVLink scale-up domain. That can reduce the need to cross an external network for workloads partitioned across more than eight GPUs.
It does not mean every model automatically runs nine times faster, or that a model with weights smaller than 20 TB will fit in practice. Runtime memory also has to accommodate KV cache, activations, temporary buffers, parallelism choices and software overhead. Delivered throughput depends on the model, precision, batch size, context length and serving stack.
Peak Tensor Core figures need similar care. NVIDIA lists GB300 NVL72 FP4 performance both with and without sparsity; mixing those modes is how apparently precise but incompatible per-GPU numbers arise. Capacity planning should keep memory, fabric and workload benchmarks separate rather than multiplying one headline FLOPS figure across products.
Which platform fits which deployment?
HGX B300 is the more modular choice when an operator wants conventional server building blocks, OEM choice and incremental cluster growth. Each node supplies eight GPUs, 2.30 TB of HBM3e and a local NVLink domain, while the cluster fabric handles communication between nodes.
GB300 NVL72 is the rack-scale choice when the workload benefits from a much larger in-rack scale-up domain and the site can support its integrated power and liquid-cooling requirements. It combines 72 GPUs, 20 TB of HBM3e, 36 Grace CPUs and rack-wide NVLink as one deployment unit.
The practical comparison is therefore not “which GPU is faster?” Both are Blackwell Ultra systems. The decision is whether the deployment needs an eight-GPU OEM server building block or a 72-GPU liquid-cooled rack with a substantially larger NVLink domain.
Sources
This article was researched and fact-checked against the following sources:
- NVIDIA HGX AI Factory: Components (docs.nvidia.com)
- NVIDIA GB300 NVL72 (nvidia.com)
- NVIDIA NVL72 AI Factory: System Hardware and Components (docs.nvidia.com)
- NVIDIA Blackwell Ultra for the Era of AI Reasoning (developer.nvidia.com)