Key takeaways

  • A GB300 NVL72 rack joins 72 Blackwell Ultra GPUs into one fifth-generation NVLink domain. That domain ends at the rack boundary.
  • ConnectX‑8 adapters carry the scale-out GPU compute traffic between racks in the cited NVIDIA Enterprise Reference Architecture, which uses Spectrum‑X Ethernet.
  • BlueField‑3 serves the separate CPU-converged fabric for storage, in-band management and other north-south traffic.
  • The physical out-of-band management network is a fourth path: 1 GbE connections to rack-management endpoints through SN2201 switches.
  • A four-rack GB300 submission reached 99% scaling efficiency in one MLPerf Inference v6.1 offline test. That result does not mean NVLink spans four racks.

The short version: four traffic paths, not three

GB300 NVL72 networking is easiest to understand by asking where a packet is going. GPU traffic inside a rack uses NVLink and NVSwitch. GPU traffic that must cross a rack boundary uses a scale-out network built around ConnectX‑8. CPU, storage and in-band management traffic uses a separate converged fabric reached through BlueField‑3. Hardware management has its own 1 GbE out-of-band network.

TrafficFabric or adapterScope
GPU-to-GPU inside one rackNVLink and NVLink SwitchOne 72-GPU rack domain
GPU compute between racksConnectX‑8 on Spectrum‑X Ethernet in the cited reference designScale-out fabric
CPU, storage and in-band managementDual-port BlueField‑3 on the CPU-converged fabricNorth-south services
BMC and switch management1 GbE through SN2201 switchesPhysically separate OOB network

That separation matters because the adapter list alone can be misleading. A compute tray contains both ConnectX‑8 and BlueField‑3, but they are not interchangeable labels for the same management network.

A GB300 NVL72 rack contains 18 compute trays with four B300 GPUs per tray, for 72 GPUs in total. Nine NVLink Switch trays and a passive copper backplane connect them. The cited architecture guide specifies 130 TB/s of aggregate NVLink bandwidth for the rack.

Lenovo’s product guide independently describes the 72 GPUs as one unified fifth-generation NVLink fabric. It also states the important boundary condition: the NVLink domain is contained within a single rack, with no inter-rack NVLink scaling.

This is the principal architectural difference from an eight-GPU DGX B300 server. DGX B300 also uses NVLink and NVSwitch internally, but its NVLink domain is limited to that server. GB300 NVL72 expands the scale-up domain from one server to a full 72-GPU rack; it does not turn multiple racks into one NVLink domain.

ConnectX‑8 handles the scale-out GPU fabric

Each GB300 compute tray in the cited design includes four ConnectX‑8 adapters. In the four-rack NVIDIA Enterprise Reference Architecture described by the source, the GPU compute fabric is Spectrum‑X Ethernet. Each 800 Gb/s ConnectX‑8 connection is split into two 400 Gb/s links distributed across two network planes.

The two-plane layout is intended to provide path redundancy, reduce single points of failure and support rail-aware load distribution. This is the east-west network used when GPU communication leaves the local NVLink rack domain.

The distinction is practical: NVLink supplies high-bandwidth scale-up inside the rack, while ConnectX‑8 supplies Ethernet scale-out between racks. Describing ConnectX‑8 as only the CPU north-south path reverses its role in this reference architecture.

BlueField‑3 serves the CPU-converged network

The same compute tray includes one dual-port BlueField‑3 B3240. The cited four-rack design connects each BlueField‑3 to two separate switches for path redundancy on a CPU-converged north-south fabric.

That converged fabric carries in-band management, storage, customer-network access, support-server traffic and CPU communication. It is separate from the Spectrum‑X GPU compute fabric even though both ultimately use Ethernet switching in this design.

For operators, this is the useful mental model: ConnectX‑8 is attached to the high-throughput GPU scale-out problem, while BlueField‑3 attaches the host and service side of the rack to the converged network.

The OOB network is physically separate

Out-of-band management should not be collapsed into the BlueField‑3 data path. The reference architecture uses two SN2201 switches per NVL72 rack and provides physically separated 1 GbE management access.

Its endpoints include the compute-tray BMC, the BlueField BMC, NVLink-switch management and other rack-management interfaces. This network remains available for hardware administration and troubleshooting without sharing the primary GPU or service fabrics.

What the DGX B300 comparison actually shows

The cited DGX B300 example has eight ConnectX‑8 compute connections at 800 Gb/s each and uses Quantum‑X800 InfiniBand with Q3400‑RA switches for its external compute fabric. Storage and management use two dual-port BlueField‑3 adapters, while a separate OOB network handles BMC and switch management.

The protocols differ from the cited GB300 Enterprise design, but the organizing principle is similar: scale-up links, scale-out compute, service traffic and hardware management are separate concerns. GB300 NVL72 changes the scale-up boundary most dramatically by putting 72 GPUs behind rack-wide NVLink.

Reading the 99% scaling result correctly

NVIDIA reported that a DeepSeek‑R1 submission scaled from one GB300 NVL72 rack to four racks, or 288 GPUs, with 99% scaling efficiency in the MLPerf Inference v6.1 offline scenario. It is a useful result for that workload and configuration, not a universal efficiency guarantee.

It also should not be attributed to sixth-generation NVLink. The same source discusses sixth-generation NVLink and a 10-times packet-rate and three-times latency comparison in its separate Vera Rubin NVL72 section. GB300 uses fifth-generation NVLink, and its four-rack result combines rack-local NVLink with scale-out networking and software orchestration.

Facilities context

Networking sits inside an unusually dense system. The cited rack specification lists 20 TB of HBM, about 135 kW nominal rack power with a 132–142 kW range and a peak near 155 kW. Lenovo describes hybrid cooling: CPUs, GPUs and NVSwitch components are liquid-cooled, while components such as OSFP modules, storage and power-distribution hardware remain air-cooled.

Those constraints are another reason to keep the fabric map precise. Cabling, switching, redundancy and management access all have to be planned around a rack that is already near the limits of conventional data-centre power and cooling practice.

Bottom line

GB300 NVL72 does not run all networking through one fabric. NVLink joins 72 GPUs inside the rack. ConnectX‑8 carries scale-out GPU traffic. BlueField‑3 connects the CPU-converged service network. A separate 1 GbE OOB layer manages the hardware. Keeping those four paths distinct is the clearest way to evaluate topology, switch counts and failure domains without overstating what rack-scale NVLink can do.

Sources

This article was researched and fact-checked against the following sources: