Four different questions wearing one word
“Interconnect” does too much work in most conversations about AI datacenter hardware. Ask which interconnect a new cluster should use and the question actually bundles four separate ones: which physical and link layer moves the bits; which transport decides what happens when a bit is lost; how many independent vendors can supply that stack; and which side of the boundary between a rack-scale scale-up domain and a building-scale scale-out fabric the technology even belongs on. InfiniBand, RDMA over Converged Ethernet (RoCE), the Ultra Ethernet Consortium’s new specification, and NVLink-class proprietary scale-up links answer those four questions differently, and much of the noise in public comparisons between them comes from collapsing four axes into one number and one supposed winner.
This article works through the four approaches on their own terms, then walks the axes that actually separate them: latency, cost and vendor structure, ecosystem and software maturity, and the deeper design choice between a fabric engineered to never drop a packet and one engineered to tolerate dropping some. None of the four is simply “faster” than the others in the way a single benchmark implies, because they are not always competing to do the same job. Where network engineers who have actually operated at scale genuinely disagree — and on the question of whether a lossless fabric is a sound foundation at all, they disagree sharply — that disagreement is characterised here rather than resolved, because resolving it from the outside would mean inventing a verdict nobody who has shipped the hardware has agreed to.
InfiniBand: lossless by construction
InfiniBand is the oldest of the four and the only one designed as a lossless fabric from its link layer upward rather than adapted to become one. The InfiniBand Trade Association, the body that has maintained the specification since 1999, describes it as “an industry standard, channel-based, switched fabric interconnect architecture for server and storage connectivity” [1]. Its current release states link speeds up to 800 Gb/s and support for tens of thousands of nodes within a single subnet, with routers extending that to effectively unlimited cluster sizes [1]. The mechanism that makes losslessness structural rather than bolted on is credit-based flow control at the link layer: a sender does not transmit a packet until the receiver has advertised buffer space to hold it, so the fabric cannot be made to drop a packet for lack of room the way a conventional best-effort network can. This is architecture, not policy — it cannot be misconfigured off the way a higher-layer flow-control feature can.
That heritage shows up in current products. NVIDIA’s Quantum-2 platform, the InfiniBand generation in wide deployment as of this writing, ships 400 Gb/s per port, up to sixty-four such ports in a single switch, and a stated packet-processing rate of more than 66.5 billion packets per second bidirectionally from one switch device; NVIDIA also claims the platform enables networks of over one million 400 Gb/s nodes across a four-switch-tier DragonFly+ topology and a fourfold improvement in MPI performance over the prior generation [7]. Those last two figures are vendor claims about a specific topology and workload, not independent measurements, and should be read that way — the honest use of a vendor datasheet is as a statement of what the maker asserts its hardware can do, not as a settled comparative fact.
InfiniBand’s adoption pattern is concentrated. Mellanox, which pioneered the technology, was acquired by NVIDIA in a deal announced in March 2019 at what NVIDIA’s own release called “a total enterprise value of approximately $6.9 billion,” with the release noting that Mellanox’s InfiniBand technology “along with its high-speed Ethernet products is now used in over half of the world’s fastest supercomputers” [8]. That acquisition put the dominant supplier of InfiniBand host adapters and switches inside the company that also builds the accelerators the fabric connects — a structural fact worth separating from any judgment about whether it is good or bad, because it cuts both ways: it lets NVIDIA co-design the GPU and the fabric as one product, and it means a buyer choosing InfiniBand today is choosing a single commercial counterparty for that layer of the stack in a way that was not true when Mellanox, Qlogic, and others competed directly for InfiniBand silicon.
Adoption data bears out both InfiniBand’s continued strength and a shift underway around it. The IBTA’s own account of the November 2024 TOP500 list reports that InfiniBand and RoCE together connect “365 systems — 73% of the world’s top supercomputers,” with InfiniBand alone powering 254 of those platforms and connecting 66 of the top 100, and driving six of the ten most energy-efficient systems on the associated Green500 list [3]. That is adoption evidence, not a performance ranking, and it should be read as exactly that: a statement about how many operators chose which fabric for a given machine, not about which fabric would win a head-to-head test on identical hardware, which nobody has published.
RoCE: InfiniBand’s transport, Ethernet’s economics
RoCE keeps InfiniBand’s transport semantics — its RDMA verbs, its message-passing and storage-protocol support — and swaps out the link and physical layers for ordinary Ethernet. The IBTA describes the original RoCE specification as designed to “provide InfiniBand Transport Services on Ethernet Networks,” preserving InfiniBand’s low latency and reduced CPU overhead “eliminated the multiple data copies inside the server” while inheriting Ethernet’s link layer and network management [2]. The consequential update was RoCEv2, published as a supplement to the InfiniBand specification in September 2014, which moved the transport onto UDP/IPv4 or UDP/IPv6 so that RDMA traffic could be routed across Layer 3 networks rather than confined to a single Ethernet broadcast domain — a change the IBTA credits with making RoCE workable for “hyperscale data center deployments” [2].
The catch is in that inheritance. Ordinary Ethernet is a best-effort network; it drops packets under congestion by design, and RDMA transport, built for a lossless substrate, degrades badly when packets are dropped. RoCEv2 deployments therefore manufacture losslessness on top of Ethernet using priority-based flow control (PFC), which pauses a link’s traffic in a given priority class rather than dropping it, combined with congestion notification that tries to slow senders down before a pause becomes necessary. Zhu and colleagues’ widely cited account of this problem, developed at Microsoft, states plainly that “PFC can lead to poor application performance due to problems like head-of-line blocking and unfairness,” and introduces DCQCN, an end-to-end congestion control scheme built on the IEEE 802.1Qau quantized congestion notification standard, reporting that it “dramatically improves throughput and fairness of RoCEv2 RDMA traffic” in a three-tier Clos testbed and noting it was implemented in Mellanox NICs and deployed in Microsoft’s own datacenters [10].
Guo and colleagues’ separate account of scaling RoCEv2 at Microsoft is the most candid public record of what that engineering actually costs in production. They report having to address “the safety challenges brought by PFC-induced deadlock (yes, it happened!), RDMA transport livelock, and the NIC PFC pause frame storm problem,” building a DSCP-based PFC scheme, configuration management and monitoring tooling, and an RDMA-specific variant of Pingmesh to catch problems before they cascaded — and conclude that “the safety and scalability issues of running RoCEv2 at scale can all be addressed” [12]. That “yes, it happened!” is not incidental colour; a real production deadlock in a network meant to be engineered against deadlock is exactly the kind of failure mode that fuels the disagreement examined later in this article.
Other operators have converged on RoCE for reasons distinct from InfiniBand’s engineering. Meta’s account of its backend RoCE fabric for distributed AI training describes a two-stage Clos topology, its own fork of NCCL, and — notably — abandoning DCQCN on its 400G deployments in favor of receiver-driven traffic admission, in which the collective communication library coordinates directly with the RoCE transport and receivers issue clear-to-send signals to cap in-flight traffic, backed by high-priority switch queuing for those control packets [13]. Alibaba’s account of its HPN network for large language model training reports a different structural choice again: a two-tier, dual-plane architecture connecting 15,000 GPUs in a single pod — a scale the authors say would traditionally require a three-tier Clos — paired with a dual top-of-rack design specifically to remove a single point of failure, reporting 14.9 percent higher training throughput than a conventional datacenter network design and no top-of-rack-related single-node failure across more than eight months in production [14]. These are not competing benchmark scores measured on comparable hardware and should not be read as a ranking; they are different operators solving the same PFC-and-congestion problem with different topology and congestion-control choices, each reported on its own infrastructure under its own methodology.
Ultra Ethernet: designing for loss instead of against it
The Ultra Ethernet Consortium’s response to the RoCE experience is not to build a better PFC. Its Specification 1.0, released 11 June 2025, covers “applications, transport protocols, congestion control, direct memory access, Ethernet link and PHY technologies, and network security” and is described by the consortium as delivering “Modern RDMA for Ethernet and IP” aimed at scaling “to millions of endpoints,” built explicitly around open standards intended to avoid vendor lock-in while remaining compatible with globally deployed Ethernet equipment [4]. The design philosophy is the meaningful departure: rather than treating packet loss and reordering as failures the network must be engineered to prevent, UEC’s transport is built to expect and recover from them directly — a modern transport layer doing selective retransmission and tolerating multipath reordering, instead of a link layer promising never to need it.
This is exactly the direction the academic critique of RoCE had pointed. Mittal and colleagues’ widely discussed paper argues that the industry’s dependence on PFC for lossless RDMA is a fragile foundation and proposes an improved RoCE NIC design — Irn — built around “simple changes to the RoCE NIC for better handling of packet losses,” implicitly making the case that a transport tolerant of loss is preferable to a link layer engineered to forbid it [11]. UEC’s Specification 1.0 is the industry-consortium-scale version of that same argument, developed independently and by a different, much larger group of vendors, but converging on the same structural conclusion: stop assuming losslessness and build the transport to survive without it.
What UEC does not yet have is deployed silicon and years of production hardening. Broadcom, a founding member of the consortium, shipped the first UEC-compliant switch silicon it has publicly detailed — Tomahawk Ultra — in July 2025, reporting 250 nanoseconds of switch latency at 51.2 Tbps of throughput and, paired with its own Scale-Up Ethernet specification, sub-400-nanosecond XPU-to-XPU latency including transit through the switch, achieved partly by cutting Ethernet header overhead from 46 bytes down to as little as 10 [16]. That is one vendor’s first-generation part, reported in trade coverage rather than an independent multi-vendor benchmark, and the specification itself is barely a year old as this is written — genuinely earlier in its maturity curve than InfiniBand’s twenty-five years or RoCE’s decade of hard-won production lessons. The gap between “a strong specification with credible first silicon” and “an ecosystem with the operational scar tissue Meta, Microsoft, and Alibaba have already accumulated with RoCE” is real, and it is the honest reason UEC cannot yet be judged on the same footing as the other two.
Proprietary scale-up: a different layer entirely
NVLink-class interconnects are not really competing in the same category as the other three, and treating them as if they were is where a surprising number of comparisons go wrong. InfiniBand, RoCE, and UEC are all scale-out fabrics: they connect nodes across a datacenter over a routed network. NVLink is a scale-up interconnect: it connects accelerators within a single node or a tightly bounded rack-scale domain over dedicated point-to-point and switched links, using a different protocol family built for memory-semantic load and store operations rather than a routed network transport. NVIDIA’s own figures illustrate the gap in kind as much as degree: fourth-generation NVLink offers 900 GB/s per GPU, fifth-generation 1,800 GB/s, and sixth-generation 3,600 GB/s — described as “over 14x the bandwidth of PCIe Gen6” — with NVLink Switch aggregate bandwidth reaching 260 TB/s across 72 GPUs wired all-to-all in the sixth generation [6]. DeepSeek’s technical report gives a concrete, independently reported ratio between the two layers on real hardware: in their H800 cluster, “NVLink offers a bandwidth of 160 GB/s, roughly 3.2 times that of IB (50 GB/s),” and their MoE routing design deliberately limits each token to at most four nodes specifically “thereby reducing IB traffic” — an explicit case of a scale-out bandwidth ceiling reaching back into model architecture [9].
That gap is exactly what has drawn Ethernet-based challengers into the scale-up domain rather than leaving it to NVLink alone. Broadcom frames Tomahawk Ultra explicitly against NVLink’s scale-up numbers, with reporting on the launch noting the chip “can scale to 1024, compared to NVLink’s 72” when deployed on Broadcom’s own Scale-Up Ethernet specification, and Broadcom representatives describing the goal as “Ethernet that works with any endpoint” against what they characterise as proprietary lock-in [16]. A separate, Ethernet-adjacent effort — the UALink Consortium — is building an open alternative aimed at the same job: its 200G 1.0 specification defines “200G per lane scale-up connection for up to 1,024 accelerators within a single AI computing pod,” pursued as “an open industry Interconnect standard” for accelerator-to-accelerator communication with direct load, store, and atomic memory-access operations [5]. Whether an open scale-up standard from a broad multi-vendor consortium can match a vertically co-designed proprietary link on latency and software maturity at the point accelerators actually ship is not yet decided by any measurement all three camps would accept as fair, and the honest position is that it is unresolved rather than settled in either direction.
One more example complicates any clean two-layer story. HPE Cray’s Slingshot, analysed in depth by De Sensi and colleagues, is “based on high-radix switches” enabling exascale and hyperscale networks in at most three switch-to-switch hops, and — notably — “uses an optimized Ethernet protocol, which allows it to be interoperable with standard Ethernet devices while providing high performance to HPC applications,” with the authors reporting that applications on Slingshot “are less affected by congestion compared to previous generation networks” through its adaptive routing and congestion control [15]. Slingshot is neither a pure open-standard Ethernet fabric nor a closed proprietary one in the NVLink sense; it sits deliberately between the categories, Ethernet-compatible at the wire but carrying proprietary routing and congestion behaviour above it. Its existence is a useful check against treating “open versus proprietary” and “scale-up versus scale-out” as the same binary — they are two separate axes, and real fabrics occupy more of the resulting grid than a four-way comparison first suggests.
Latency: real numbers, measured differently
Every latency figure cited in the reporting above was measured on specific hardware, at a specific link speed, sometimes under a specific topology, by an entity with a commercial interest in the result. That does not make the figures worthless; it makes them non-comparable without matching the conditions, which none of the public disclosures do for each other. Broadcom’s reported 250 nanoseconds of switch latency at 51.2 Tbps for Tomahawk Ultra and its sub-400-nanosecond claimed XPU-to-XPU figure under its own Scale-Up Ethernet stack are switch-level and link-level numbers for a specific first-generation part [16]. NVIDIA’s Quantum-2 material publishes packet-processing throughput and topology-scale claims but does not publish a comparable port-to-port latency figure in the material examined here [7]. Zhu and colleagues frame the target application requirement for RoCEv2 deployments at Microsoft as “ultra-low latency (< 10 μs per hop)” — an application-level design target for a deployed fleet, not a hardware datasheet number, and roughly an order of magnitude coarser than the switch-level nanosecond figures quoted for newer silicon, because it is measuring a different point in the stack under real production load rather than a switch ASIC in isolation [10].
This is the concrete version of the instruction that runs through this whole comparison: a nanosecond switch-latency claim, a microsecond application-level target, and a “4X MPI performance” vendor claim over a prior generation are three different kinds of number, measuring three different things, none of them substitutable for the others, and none of them licensing a claim that one fabric is simply “faster” than another in general. Where a genuinely comparable number exists — DeepSeek’s 3.2x NVLink-to-InfiniBand bandwidth ratio, measured on their own cluster and reported in their own technical report — it says something real about that specific pairing on that specific hardware, and nothing directly transferable to a RoCE or UEC deployment on different silicon [9].
Cost and ecosystem: what a fabric buys beyond its datasheet
Cost differences between these fabrics show up structurally rather than in a public price list, and the structural story is legible from vendor concentration and consortium membership even without a dollar figure. InfiniBand’s silicon supply is now effectively a single-vendor question following NVIDIA’s 2019 acquisition of Mellanox [8], which the IBTA’s own TOP500 reporting shows has not slowed adoption — InfiniBand still powers 254 of the November 2024 TOP500 systems — but which does mean a buyer choosing InfiniBand is choosing one commercial counterparty for host adapters and switches at once [3]. RoCE runs on Ethernet silicon supplied by a genuinely competitive field — Broadcom, NVIDIA’s own Spectrum line, Cisco, Arista, and others all ship RoCE-capable Ethernet switching — and that IBTA report’s own headline finding, that RoCE reached the TOP100 for the first time “in nearly a decade” at the November 2024 list, is itself adoption evidence for that broader ecosystem gaining ground on the merits its operators cared about, not a claim that RoCE now outperforms InfiniBand [3].
UALink and UEC are both explicit bets that vendor breadth is worth the maturity gap it currently carries. UALink was formed with backing from a broad set of hyperscalers and silicon vendors pursuing “an open industry Interconnect standard” specifically so that no single company controls the scale-up layer the way NVIDIA effectively does today [5], and Broadcom’s own framing of its Ethernet-based scale-up push is explicit about the same motive, describing hyperscaler demand for “fungible interfaces on their GPU” rather than a single proprietary connector [16]. Ecosystem maturity, though, is not the same claim as vendor breadth — it also means software: InfiniBand’s verbs API and subnet manager tooling, and RoCE’s shared inheritance of that same verbs layer, have a decade or more of production debugging behind them across multiple hyperscalers, documented in detail by Guo and colleagues’ own account of the monitoring and incident-response tooling Microsoft had to build around RoCE in production [12]. UEC and UALink do not yet have an equivalent public record, because they do not yet have an equivalent number of years in large-scale production; that is a fact about timing, not a verdict on the designs.
Where the disagreement is real
The sharpest disagreement among the sources examined here is not about a number. It is about whether a lossless fabric — InfiniBand’s approach, and RoCE’s attempt to reconstruct it on Ethernet — is the right foundation at all. Mittal and colleagues’ position, developed from outside any single vendor’s production fleet, is that dependence on PFC for losslessness is a structural liability worth re-architecting around, because PFC’s failure modes are severe and hard to eliminate rather than merely inconvenient [11]. Guo and colleagues’ position, developed from years of actually running RoCEv2 at Microsoft’s scale, is that the specific failure modes — including a production PFC-induced deadlock they candidly flag with “yes, it happened!” — were each addressed, and that “the safety and scalability issues of running RoCEv2 at scale can all be addressed” [12]. Meta’s own operational choice sits in between: rather than trusting DCQCN’s standard congestion-control path on its newest 400G fabric, the company disabled it and built a different, receiver-driven admission scheme instead [13] — which is neither a vindication of PFC-based losslessness as designed nor an abandonment of it, but a third position: keep the lossless link layer, replace the piece of the congestion-control stack that did not hold up in practice.
The Ultra Ethernet Consortium’s Specification 1.0 is best read as a bet that the Mittal position generalises: that the industry should stop trying to perfect a lossless Ethernet fabric and instead build a transport that does not need one [4]. Whether that bet pays off depends on questions no public disclosure yet answers cleanly — whether a loss-tolerant transport can match a well-tuned PFC deployment’s tail latency under the bursty, low-entropy traffic patterns AI training produces, and whether UEC silicon reaches the same production hardening RoCE has accumulated over a decade, faster than RoCE’s remaining rough edges get sanded down further. Both are open questions, not foregone conclusions in either direction, and treating either outcome as already decided would be exactly the invented ranking this article has tried to avoid.
What follows for someone actually choosing
A few practical distinctions survive everything above without needing a benchmark.
Match the fabric to the boundary, not the brand. NVLink-class links belong inside the scale-up domain regardless of which scale-out fabric sits outside it; the choice between InfiniBand, RoCE, and UEC is a scale-out question and should be evaluated as one, separately from the scale-up choice.
Treat vendor latency and throughput figures as claims about specific hardware under specific conditions, not as comparable data points, until two vendors publish numbers measured the same way on the same workload — which, in the sources examined here, none have.
Weigh ecosystem maturity as a real, separate cost from raw specification quality. RoCE’s production scar tissue at Microsoft, Meta, and Alibaba is not evidence that RoCE is technically superior to UEC; it is evidence that RoCE has had a decade longer to accumulate the operational tooling a fabric needs at scale.
Expect the “open versus proprietary” and “scale-up versus scale-out” axes to keep producing hybrids like Slingshot, rather than resolving into two clean camps — a fabric can be Ethernet-compatible at the wire and proprietary in its routing and congestion behaviour at once.
Predictions, with what would falsify them
These are forecasts, kept separate from the sourced comparison above. Horizon: 15 August 2029.
One. UEC-compliant silicon will appear from at least three independent switch or NIC vendors in production AI clusters, not only in vendor demonstrations. Observable indicator: operator disclosures analogous to Meta’s and Alibaba’s RoCE accounts, naming UEC in production. Disconfirmed if by 2029 UEC deployments remain confined to lab and pilot environments across the industry.
Two. The InfiniBand-versus-RoCE adoption split will continue to move toward RoCE in TOP500-class disclosures without InfiniBand’s absolute deployment count falling. Observable indicator: IBTA’s own future TOP500 reports. Disconfirmed if RoCE’s TOP100 share reverses back toward its pre-2024 near-absence.
Three. At least one open scale-up standard — UALink or a Scale-Up Ethernet variant — will ship in a publicly disclosed accelerator cluster at a scale exceeding 512 accelerators in a single domain. Disconfirmed if by 2029 all publicly disclosed scale-up deployments above that size remain NVLink-based.
Four. Public operator accounts will increasingly report congestion-control failures and fixes as engineering narratives — the way Guo and colleagues and Meta have — rather than as vendor-reported benchmark wins, because that is the only currency other operators have shown they trust. Disconfirmed if future comparisons revert to unqualified vendor throughput claims as the dominant form of public evidence.
None of these requires a technology to fail. They follow from the pattern already visible: adoption evidence, operational disclosure, and consortium formation are the leading indicators in this space, well ahead of any independent cross-vendor benchmark that does not yet exist.
What to take away
InfiniBand, RoCE, Ultra Ethernet, and NVLink-class scale-up links are not four attempts at the same product. InfiniBand is a fabric engineered to be lossless from its link layer up, now supplied by one dominant vendor with a quarter-century of production hardening behind it. RoCE keeps InfiniBand’s transport but runs it over Ethernet’s more competitive, more familiar hardware ecosystem, at the cost of having to manufacture losslessness through priority flow control and congestion notification that operators including Microsoft and Meta have each had to debug and, in Meta’s case, partly replace in production. Ultra Ethernet is a bet, not yet fully tested at scale, that the entire premise of manufactured losslessness should be abandoned in favor of a transport built to tolerate loss directly — a position the academic RDMA-networking literature had already been arguing for years before the consortium formalised it. NVLink-class interconnects are not competing with any of the three; they solve a different problem, inside a boundary the other three do not cross, and the challenges now aimed at that boundary from UALink and Scale-Up Ethernet are themselves unresolved bets rather than settled outcomes.
The comparison that matters is not which of the four wins. It is which axis — latency at a specific point in the stack, cost and vendor structure, ecosystem maturity measured in years of production debugging, or a considered position on whether losslessness belongs in the link layer or the transport — actually decides the engineering question in front of you, and refusing to let a vendor’s headline number answer a question it was never measured against.