SemiconductorsFEATURE

NVLink, UALink, and Ultra Ethernet: The Bandwidth Battle Over AI Chip Scale-Up Interconnects

N
NathanTechnology Editor · Technical Lead
Published · Updated
Sixth-gen NVLink reaches 260 TB/s in a 72-GPU domain; the open camp answers with UALink (scale-up) and Ultra Ethernet (scale-out).

What Is the Interconnect Hierarchy in AI GPU Clusters, and Where Does the Bandwidth Bottleneck Bite?

NVIDIA divides AI GPU cluster interconnects into two layers: scale-up within a compute domain and scale-out across data-center serversCITE:E1. In NVIDIA's framing, scale-out networks connect servers across the data center, while scale-up networks let the GPUs inside a single domain behave as one engine of computeCITE:E1. NVIDIA positions NVLink as "the purpose-built scale-up networking fabric for AI factories"CITE:E1.

NVIDIA states that this layer is not a peripheral detail — all-to-all bandwidth and latency are critical to AI factory performanceCITE:E6. The company specifically warns that if experts in a model sit behind a low-bandwidth or high-latency fabric, the gains from expert parallelism can be erased by communication overheadCITE:E6. That framing is why scale-up bandwidth, not just raw compute, has become a competitive front.

What Bandwidth Does NVLink Deliver, from GPU-to-GPU Links to the GB200 NVL72 Rack?

NVIDIA's sixth-generation NVLink provides 3.6 TB/s of bidirectional GPU-to-GPU bandwidth per GPU, scaling to 260 TB/s of rack-level bandwidth across a 72-GPU domainCITE:E2.

At the system level, the NVIDIA GB200 NVL72 connects 36 Grace CPUs and 72 Blackwell GPUs in a rack-scale, liquid-cooled designCITE:E3. NVIDIA states its NVLink Switch System provides 130 TB/s of low-latency GPU communication for AI and high-performance computing workloads within that rackCITE:E3.

Building blockEntityScopeKey figure
NVLink (6th gen)NVIDIAPer-GPU, bidirectional3.6 TB/s
NVLink Switch SystemNVIDIA72-GPU domain260 TB/s
GB200 NVL72 rackNVIDIA36 Grace CPUs + 72 Blackwell GPUs130 TB/s
UALink 1.0 (200G)UALink ConsortiumUp to 1,024 accelerators/pod200G per lane

What Do the UALink and Ultra Ethernet Open Standards Specify, and How Do Their Roles Differ?

UALink Consortium's 200G 1.0 specification supports up to 1,024 accelerators within a single AI computing pod, while Ultra Ethernet Consortium addresses a separate layer of the networkCITE:E4CITE:E5.

UALink Consortium describes UALink as "the open scale-up interconnect for next generation AI workloads"CITE:E4. Its UALink 1.0 Specification, published April 8, 2025, "enables 200G per lane scale-up connection for up to 1,024 accelerators within an AI computing pod"CITE:E4. That positions UALink as a direct open alternative to NVLink's scale-up role described above.

Ultra Ethernet Consortium, by contrast, states its goal is to "deliver an Ethernet based open, interoperable, high performance, full-communications stack architecture to meet the growing network demands of AI & HPC at scale"CITE:E5. The consortium explicitly defines "AI at scale" as the "GPU scale-out network"CITE:E5 — the layer NVIDIA itself separates from scale-upCITE:E1. In other words, UALink competes with NVLink at the scale-up layer, while Ultra Ethernet occupies the scale-out layer alongside it rather than against it.

Who Is Lining Up Behind the Open Interconnect Camp Against NVIDIA's NVLink?

UALink Consortium counts more than 85 member companies backing its open scale-up standardCITE:E7. The consortium states its board includes Alibaba, AMD, Apple, Astera Labs, AWS, Cisco, Google, HPE, Intel, Meta, Microsoft, and SynopsysCITE:E7.

That roster spans cloud operators, chipmakers, and networking vendors, positioned around a single open industry standard defined by the UALink ConsortiumCITE:E7.

What this means: NVIDIA has already shipped a scale-up fabric — sixth-generation NVLink at 3.6 TB/s per GPU and 260 TB/s per domain, running inside the GB200 NVL72 at 130 TB/sCITE:E2CITE:E3 — for the exact layer it identifies as critical to AI factory performanceCITE:E6. UALink Consortium's published 200G-per-lane, 1,024-accelerator specification targets that same scale-up layerCITE:E4, and it carries the backing of 85-plus companies including several of NVIDIA's largest customers and rivalsCITE:E7. Ultra Ethernet Consortium's open stack, meanwhile, is defined against the adjacent scale-out layerCITE:E5, so it does not directly contest NVLink's scale-up bandwidth figures — it sits beside that fight rather than inside it.

📊 Evidence

FAQ

What Is the Interconnect Hierarchy in AI GPU Clusters, and Where Does the Bandwidth Bottleneck Bite?

NVIDIA divides AI GPU cluster interconnects into two layers: scale-up within a compute domain and scale-out across data-center serversCITE:E1.

What Bandwidth Does NVLink Deliver, from GPU-to-GPU Links to the GB200 NVL72 Rack?

NVIDIA's sixth-generation NVLink provides 3.

What Do the UALink and Ultra Ethernet Open Standards Specify, and How Do Their Roles Differ?

UALink Consortium's 200G 1.0 specification supports up to 1,024 accelerators within a single AI computing pod, while Ultra Ethernet Consortium addresses a separ…

Who Is Lining Up Behind the Open Interconnect Camp Against NVIDIA's NVLink?

UALink Consortium counts more than 85 member companies backing its open scale-up standardCITE:E7.

📎 Sources

  1. developer.nvidia.com
  2. nvidia.com
  3. ualinkconsortium.org
  4. ultraethernet.org
Author's TakeNathan

The numbers show why this is a two-front contest, not a single race. NVLink's 3.6 TB/s per-GPU bandwidth and 260 TB/s domain-level figure are already running in production inside the GB200 NVL72, which delivers 130 TB/s of low-latency GPU communication at rack scale. UALink's 200G-per-lane, 1,024-accelerator specification, backed by a board of 85-plus companies including AMD, AWS, Google, Intel, and Microsoft, is by comparison still a published standard rather than a shipping system. Ultra Ethernet is not a direct NVLink rival at all — its own consortium frames its job as the scale-out layer, not scale-up, so it competes alongside UALink and NVLink rather than against them. The metric worth watching is whether UALink's member companies ship silicon that actually hits the 200G-per-lane, 1,024-accelerator spec in real clusters, since NVIDIA's own framing of all-to-all bandwidth and latency as critical to AI factory performance is the bar any open alternative has to clear.

N
NathanTechnology Editor · Technical Lead

Related

FEATURE

Why Solar-Plus-Storage Is the Most Practical Power Fix for AI Data Centers Right Now

Utility-scale solar has become one of the cheapest new power sources, with costs down roughly 90% since 2010, while 2024 deployment volume outpaced every other generation technology. Paired with record-low battery prices, solar-plus-storage is already powering an AI data center in Arizona — though it still falls short of full 24-hour dispatchable baseload.

Nathan ·
FEATURE

2026's AI Enforcement Collision: EU Fines Activate as California Tightens and Washington Pushes Back

2026 marks the European Union AI Act's real enforcement start: the AI Office gains fining power over general-purpose AI on August 2, while high-risk system deadlines are pushed to 2027 and 2028. California activates two new laws on January 1 covering frontier-developer safety disclosure and training-data transparency. The federal government moves the opposite direction, ordering a Justice Department task force to challenge state AI laws, naming California's SB 53 as a target.

EffectStory 編輯部 ·
FEATURE

Who's Actually Flying Air Taxis? China's EHang Carries Passengers While US Rivals Chase Certification and Europe's Two Pioneers Collapse

China's EHang (億航) is the only eVTOL maker actually flying paying passengers today, holding a full Chinese type, production, and operating certificate set and delivering 221 aircraft in 2025. US rivals Joby and Archer remain in certification or pre-launch stages, while Germany's Lilium and Volocopter both went insolvent in 2024–2025, with Volocopter absorbed by a Chinese buyer.

EffectStory 編輯部 ·