Robots Atlas>ROBOTS ATLAS
Infrastructure

IB

2000ActivePublished: 8 May 2026Updated: 8 May 2026Published
Key innovation
A switched-fabric network with native RDMA, lossless credit-based link-level flow control, and sub-microsecond latencies — designed from the ground up as an HPC/AI interconnect rather than as a bolt-on over an existing stack.
Category
Infrastructure
Abstraction level
Pattern
Operation level
TrainingDeploymentSystem
Use cases
LLM training clusters (NVIDIA SuperPOD, DGX, frontier compute)TOP500 supercomputersScale-out storage (Lustre, GPFS, NVMe-oF)Oracle Exadata databasesScientific simulations / CFD / climate

How it works

Each host carries a Host Channel Adapter (HCA) — an intelligent NIC that implements the entire protocol stack in hardware. The application uses the verbs API (ibv_post_send) to post an RDMA WRITE/READ/SEND or an atomic operation; the HCA directly reads/writes remote memory with zero copies and no CPU involvement on the remote side. The switched fabric uses a Subnet Manager to compute paths (linear forwarding tables) and credit-based flow control: a sender transmits only when the receiver has buffer credit available, guaranteeing losslessness. Physical layer: links are aggregated (1×/4×/8×/12×) with QSFP (up to HDR) and OSFP (NDR and beyond) connectors, copper up to 10 m, fiber up to 10 km.

Problem solved

Traditional Ethernet-with-TCP/IP introduced high latency, CPU overhead, and lossy behavior that disqualified it as an HPC/AI interconnect. InfiniBand solves this with native RDMA, lossless link-level flow control, and a switched-fabric topology from layer 1 up.

Components

Host Channel Adapter (HCA)Host hardware endpoint

Host-side adapter that implements the IB transport stack in hardware and serves the RDMA verbs (send, receive, write, read, atomic).

IB SwitchForwarding plane

Fabric switch that forwards IB packets between HCAs based on the linear forwarding table installed by the Subnet Manager.

Subnet Manager (SM)Control plane

Control-plane component (typically run on one of the nodes or in a switch) that discovers topology, assigns LIDs, and programs routing tables in switches.

Official

Verbs APISoftware interface

IBTA-standardized set of programming operations (ibv_post_send, ibv_open_device, ibv_reg_mr…) implemented by the libibverbs library (OFED).

Implementation

Implementation pitfalls
Vendor lock-in (NVIDIA/Mellanox)High

After the Mellanox acquisition (2019) and Intel's exit (Omni-Path), NVIDIA is effectively the sole IB hardware vendor.

Fix:Choose RoCE / Ethernet as the alternative, or pursue a multi-vendor strategy using Ultra Ethernet.
Subnet Manager as single point of failureMedium

Master SM failure blocks new path setup; a standby SM must be configured.

Fix:Master/standby SM, monitoring, and automatic failover.
No native IP routingMedium

IB is a dedicated fabric — IPoIB or a gateway is required to interoperate with the broader IP infrastructure.

Fix:IPoIB, EoIB, gateway switches.
CapEx / OpEx costMedium

IB hardware (HCAs, switches, cabling) is typically more expensive than equivalent Ethernet at the same line rate.

Evolution

1999
IBTA founded (merger of NGIO and Future I/O)

NGIO (Intel) and Future I/O (Compaq, IBM, HP) merge into the InfiniBand Trade Association.

2000
InfiniBand Architecture Specification 1.0
Inflection point

First release of the IB architecture specification.

2001
Mellanox InfiniBridge — first 10 Gbit/s product

Mellanox ships the first commercial InfiniBand products at 10 Gbit/s line rate (SDR).

2005
InfiniBand lands in Linux Kernel 2.6.11

OpenIB Alliance (later OpenFabrics) integrates the IB stack into the mainline kernel.

2014
IB becomes the most-used TOP500 interconnect
Inflection point

After years of HPC ecosystem growth, InfiniBand becomes the dominant interconnect on the TOP500 list.

2019
NVIDIA acquires Mellanox for USD 6.9 billion
Inflection point

The acquisition makes IB a strategic component of NVIDIA's AI platform — the Quantum (switches) and ConnectX (HCA) lines.

2022
NDR — 400 Gbit/s

Introduction of NDR (Quantum-2, ConnectX-7) — the scale-out fabric of frontier-class AI clusters.

2024
XDR — 800 Gbit/s (Quantum-X800)
Inflection point

NVIDIA announces Quantum-X800 and ConnectX-8 as the next-gen fabric for Blackwell GPUs.

Hyperparameters (configurable axes)

Data rate (SDR/DDR/QDR/FDR/EDR/HDR/NDR/XDR)Critical

Per-lane bandwidth generation — from 2.5 Gbit/s (SDR) up to 200 Gbit/s (XDR).

EDR (100 Gbit/s 4×, 2014)
HDR (200 Gbit/s 4×, 2018)
NDR (400 Gbit/s 4×, 2022)
XDR (800 Gbit/s 4×, 2024)
Lane width (1×/4×/8×/12×)High

Number of aggregated physical lanes per port. 4× is the standard; 12× is used switch-to-switch.

Fabric topologyCritical

Fat tree, dragonfly, torus — affects bisection bandwidth, cost, and diameter.

MTUMedium

IB packet size — typically 256 B to 4 KB (max).

Parallelism

Parallelism level
Fully parallel

A switched fabric with multi-rail HCAs and adaptive routing enables parallel communication between thousands of GPUs without a single-link bottleneck.

Scope
TrainingInferenceAcross devices

Hardware requirements

Primary

IB is the primary scale-out fabric of the NVIDIA DGX/SuperPOD platforms for H100/H200/B200 GPU clusters.