A network before there was a market for one
Long before “datacenter interconnect” was a category anyone needed a name for, Robert Metcalfe co-invented Ethernet in 1973 while working at Xerox’s Palo Alto Research Center, building a network to connect PARC’s experimental computers to a fast new laser printer [1]. That origin matters for a reason beyond trivia: Ethernet was built to be a general-purpose, best-effort broadcast medium for an office building, not a fabric for thousands of tightly coupled processes exchanging small messages under a latency budget measured in microseconds. For roughly the first quarter-century of its existence that mismatch did not matter, because nothing needed the second thing badly enough to pay for it.
That changed as high-performance computing moved from monolithic vector supercomputers toward clusters of many smaller, commodity nodes through the 1990s. A single vector machine’s internal backplane had no equivalent outside its own cabinet; once the same aggregate performance had to be assembled from hundreds or thousands of separate, off-the-shelf nodes instead, the wire joining them stopped being an afterthought and became the thing that decided whether the cluster behaved like one coherent machine or like a loose collection of computers occasionally exchanging files. A cluster is only as fast as the network that lets its nodes act like one machine, and the Ethernet of that decade — shared media, best-effort delivery, no native concept of remote memory access — was not built for that job. The history that follows is the story of what filled the gap, what Ethernet’s own camp did in response, what happened once the accelerator rather than the node became the unit of compute, and what a fourth attempt at an answer looks like now that all of that has collided with AI-scale training.
Three lineages run through this account rather than one. A fabric purpose-built for clusters arrived first. Ethernet’s camp spent the following decade and a half teaching it the one capability it had always lacked. Then, as accelerators inside a single chassis became the tightest-coupled compute of all, a third lineage appeared that never leaves the node. Each answered the bottleneck in front of it at the time it appeared; none of the three has fully retired the others.
1999-2000: two rejected proposals merge into one specification
By the late 1990s, two rival vendor camps were each proposing a next-generation I/O interconnect to replace the shared parallel bus inside and between servers. Compaq, IBM and Hewlett-Packard were backing a proposal called Future I/O. Intel was leading a competing effort called Next Generation I/O, with Microsoft and Sun Microsystems among its backers. Neither proposal defeated the other in the market. Instead, on August 27, 1999, the two camps merged their efforts into a single organization, the InfiniBand Trade Association [2]. That founding-by-merger, rather than founding-by-victory, is itself part of the pattern this history keeps repeating: a new fabric usually arrives as a truce between rival vendors racing toward the same bottleneck, not as one company’s unilateral invention.
Just over a year after that merger, the association published what the merger had been organized to produce. The InfiniBand Architecture Specification, Release 1.0, was finalized on October 24, 2000, issued in two volumes — general specifications and physical specifications — from Portland, Oregon [3]. The specification defined a switched, point-to-point fabric with credit-based flow control built in from its first release, rather than retrofitted onto a shared-medium design as an afterthought, and native support for remote direct memory access: a receiving node’s application memory could be written to directly from across the fabric, without the sending and receiving CPUs copying the data through their own kernel network stacks along the way. That RDMA semantics, present from Release 1.0 rather than added later, is the specific capability the rest of this history keeps circling back to, because it is the one thing ordinary Ethernet of that era categorically could not do.
Ethernet is taught not to drop packets
RDMA’s central promise — write directly into a remote application’s memory without an intervening copy — depends on the network not silently discarding the packet carrying that write. InfiniBand had credit-based flow control designed into it from the outset, so a sender never transmitted faster than a receiver had buffer space to absorb. Ordinary Ethernet had no equivalent mechanism at the priority level; it could pause an entire link under IEEE 802.3’s older PAUSE frame, but not selectively protect one class of traffic while leaving everything else on the same cable running at full, loss-tolerant speed. Making RDMA viable on Ethernet’s chassis and cabling economics meant answering that gap directly, and the industry answered it in two separate, coordinated pieces roughly a decade after InfiniBand’s own specification had shipped.
The InfiniBand Trade Association itself supplied the transport-layer half first. In 2010 it released the original RDMA over Converged Ethernet specification, carrying InfiniBand’s low-latency, low-CPU-overhead RDMA semantics onto an ordinary Ethernet frame, and describing the result as delivering high network utilization with support for message-passing, socket and storage protocols over the same physical fabric [4]. That original RoCE, however, was confined to a single Ethernet broadcast domain: its packets carried no routable header, so a RoCE flow could not cross a Layer 3 boundary onto a different subnet.
The link-layer half arrived from the IEEE 802.1 working group. Its priority-based flow control project received authorization on March 27, 2008, and the finished standard, IEEE 802.1Qbb-2011, was approved on June 16, 2011 [5]. Where the older PAUSE mechanism halted an entire link, priority-based flow control operates per traffic class as identified by the VLAN priority tag, so a single physical port and cable can carry a lossless, RDMA-bearing class of traffic alongside ordinary best-effort traffic at the same time, pausing only the class that needs it [5]. That is the specific engineering move that let RoCE run on commodity Ethernet switches without dedicating the whole link to lossless behavior.
Neither piece was useful without more raw bandwidth to spend on it, and that arrived from a separate direction. Approving IEEE 802.3ba-2010 at its June 2010 Standards Board meeting, the IEEE ratified the first Ethernet amendment to specify two new speeds simultaneously — 40 and 100 gigabits per second — with physical-layer specifications spanning backplane, twinax copper, multimode and singlemode fiber [6]. The ten-gigabit ceiling that had defined commodity Ethernet through most of the 2000s stopped being the trade-off a cluster operator had to accept in exchange for RDMA.
The fabric earns its place: InfiniBand’s climb through the Top500
While Ethernet’s camp was rebuilding lossless behavior and bandwidth from separate directions, InfiniBand’s own fabric was compounding a lead in the specific machine rooms that cared most about interconnect latency and message rate: high-performance computing clusters ranked on the twice-yearly Top500 list of the world’s fastest measured systems. The InfiniBand Trade Association’s own account of the June 2012 list describes an inflection point in those terms: “InfiniBand’s presence increased on the list to become the most used interconnect for the first time,” connecting 25 of the 30 most compute-efficient systems on the list including the top two, with InfiniBand-based system performance reported to have grown 69 percent between June 2011 and June 2012 alone, and the newer Fourteen Data Rate generation increasing the number of systems it connected roughly tenfold over the year before [7]. That characterization comes from the technology’s own trade association, describing its members’ standing, and should be read as an interested party’s framing of a favorable result rather than independent analysis. What makes the underlying figure checkable rather than merely promotional is that Top500 itself is an externally compiled, standing census of measured systems, published on a fixed public schedule that the association does not control.
Two years after that milestone, the association returned to close RoCE’s most significant limitation. RoCEv2, announced on September 16, 2014, extended the RDMA transport across Layer 3 by encapsulating it in standard UDP over IPv4 or IPv6, so a RoCE flow could now be forwarded, load-balanced, monitored, metered and firewalled using ordinary IP-routing infrastructure rather than being confined to a single Ethernet broadcast domain — with the specification credited to contributions from at least ten member companies and endorsed publicly by IBM, Microsoft, Mellanox and Emulex among others [8]. With that change, a converged-Ethernet fabric could, in principle, do what InfiniBand had been purpose-built to do from Release 1.0: carry RDMA traffic across a routed, multi-subnet datacenter network rather than a single flat segment.
Accelerators create a third lineage
Everything so far describes the fabric connecting one node to the next. A separate problem was forming inside the node itself, as GPUs stopped being peripherals attached over a general-purpose expansion bus and became the primary compute resource that needed to talk directly to its neighbors at far higher bandwidth than any external network link could plausibly carry to every port in a rack.
NVIDIA’s answer arrived with the Pascal architecture. Its developer blog, published April 5, 2016 and describing the Tesla P100 accelerator, specifies NVLink as “NVIDIA’s new high-speed interconnect technology for GPU-accelerated computing,” with a single link supporting up to 40 gigabytes per second of bidirectional bandwidth and the P100 implementation supporting up to four links for an aggregate maximum of 160 gigabytes per second of bidirectional bandwidth between GPUs, enabling a program running on one GPU to execute directly against data resident in another GPU’s memory, including atomic memory operations on a remote GPU’s addresses [9]. That is a fundamentally different physical layer from anything running between racks: a point-to-point, in-node connection with no switch, no routing and no addressing scheme borrowed from either InfiniBand or Ethernet.
Two years later NVIDIA generalized the idea from a point-to-point link into a switched fabric confined to a single chassis. NVSwitch, described in an NVIDIA developer blog published August 21, 2018 in the context of the DGX-2 system, is built around an 18-port fully connected crossbar delivering 51.5 gigabytes per second per port for 928 gigabytes per second of aggregate bidirectional bandwidth in a single switch chip, connecting sixteen Tesla V100 32GB GPUs so that, in NVIDIA’s own description, they behave as “One Gigantic GPU,” with the company reporting 2 to 2.7 times faster performance on HPC and AI training benchmarks compared to two interconnected DGX-1 servers [10]. Those specific multiplier figures are the vendor’s own benchmark comparison and should be read as a vendor’s claim rather than an independently replicated result; the underlying architectural fact — a switched fabric whose entire domain lives inside one chassis, carrying more aggregate bandwidth than the external network links leaving that same chassis by a wide margin — does not depend on that comparison being taken at face value.
The consequence for this history is structural rather than a matter of one link beating another. A third lineage now existed that neither InfiniBand nor Ethernet’s converged variants were designed to compete with, because it was never meant to leave the box. Nothing about NVLink or NVSwitch required InfiniBand or Ethernet to change; the two older fabrics kept doing exactly what they had always done, one hop further out, connecting one chassis’s NVSwitch domain to the next chassis’s rather than one GPU to the next. The boundary between what runs on the in-node fabric and what runs on the external cluster network became, from this point forward, a design decision every large accelerator system had to make explicitly, and it is a boundary with no equivalent anywhere earlier in this history: InfiniBand and converged Ethernet had spent two decades arguing over the same external layer, while this new lineage simply did not compete on that layer at all.
2023: an industry consortium answers the accelerator era
By the early 2020s, training the largest AI models meant coordinating tens of thousands of accelerators across the external, scale-out network — precisely the layer that NVLink and NVSwitch had deliberately left untouched, and that InfiniBand and converged Ethernet had spent two decades building toward, but for a workload of a different character than either had originally been designed around. On July 19, 2023, nine companies — AMD, Arista, Broadcom, Cisco, Eviden (an Atos business), HPE, Intel, Meta and Microsoft — announced the formation of the Ultra Ethernet Consortium, a Joint Development Foundation project hosted by the Linux Foundation, organized to build a complete Ethernet-based communication stack architecture addressing AI and high-performance-computing workloads that the consortium’s own announcement describes as requiring “best-in-class functionality, performance, interoperability and cost-effectiveness” [11]. The founding members seeded four working groups — physical layer, link layer, transport layer and software layer — covering the stack end to end rather than one layer at a time, and the consortium opened its doors to additional member applications beginning in the fourth quarter of 2023 [11].
Read against everything before it, the Ultra Ethernet Consortium is less a new idea than a more organized version of the same move Ethernet’s camp made in 2010 and 2011: rather than a single vendor extending RoCE piecemeal or a single working group revising one IEEE amendment at a time, an entire industry coalition formed explicitly, and simultaneously, around the admission that neither InfiniBand’s decades-old fabric nor converged Ethernet’s decade-old patchwork of RoCE, RoCEv2 and priority-based flow control was, by itself, the complete answer to coordinating an AI-training cluster at the scale 2023’s frontier models already required.
Where the two older lineages stand today
The Ultra Ethernet Consortium’s formation did not settle the question of which existing fabric AI training clusters actually run on; it opened a new attempt to answer it. The InfiniBand Trade Association’s account of the November 2024 Top500 list gives the clearest available snapshot of where the two established lineages stood at that moment: 365 systems, 73 percent of the ranked list, ran on InfiniBand or RoCE, split between 254 systems on InfiniBand and 111 on RoCE-based Ethernet, with InfiniBand connecting 66 of the top 100 systems and powering six of the ten most energy-efficient systems on the associated Green500 list [13]. The same account records RoCE’s own milestone for that list: “for the first time in nearly a decade,” Ethernet-based RoCE systems re-entered the Top100, with two RoCE deployments placing 34th and 36th [13].
Read carefully, that is neither a story of InfiniBand losing ground nor of converged Ethernet definitively arriving. It is a snapshot of two fabrics coexisting within the same ranked list twelve years after InfiniBand first became the most-used interconnect on it, each retaining a distinct share rather than one displacing the other outright — precisely the kind of standing coexistence a genuinely new, third attempt at a specification would be entering into, not a settled market it could simply inherit.
2025: the specification arrives
Two years after its founding announcement, the Ultra Ethernet Consortium delivered what its founding working groups had been organized to produce. Specification 1.0, released on June 11, 2025, addresses transport protocols, congestion control, direct memory access, Ethernet link and physical-layer technologies, and network security together, and the consortium describes it as providing “Modern RDMA for Ethernet and IP” intended for “intelligent, low-latency transport for high-throughput environments” [12]. The consortium’s own scalability claim for the specification is explicit: “End-to-End Scalability — From routing and provisioning to operations and testing, UEC scales to millions of endpoints” [12]. That is the consortium’s characterization of what its own specification is designed to achieve, not a report of deployed silicon meeting that figure in production; as of this account, Specification 1.0 is a published document rather than a body of measured, independently reported deployment results, and the distinction between the two is exactly the one this history has had occasion to draw at nearly every turn since 1999.
What this history actually shows
Read end to end, the pattern is more specific than “networks keep getting faster.” Each of the fabrics in this account was a dated, named response to a bottleneck that a prior fabric was not built to carry: a merger of rejected proposals produced a cluster fabric with RDMA built in from Release 1.0 in 2000, because nothing general-purpose could do that yet; Ethernet’s own camp spent 2010 through 2014 adding, piece by piece, the specific lossless and routable behavior that cluster fabric had shipped with a decade earlier; NVLink and NVSwitch answered a bottleneck neither older lineage was ever positioned to solve, because it lived entirely inside a chassis those older fabrics never reached; and the Ultra Ethernet Consortium organized, in 2023, an entire industry’s attempt to do at coalition scale and in one coordinated specification what earlier Ethernet extensions had done piecemeal.
None of the earlier fabrics disappeared when the next one arrived. InfiniBand did not retire when RoCEv2 shipped in 2014; it still connected two-thirds of the InfiniBand-or-RoCE share of the Top500 a decade later. NVLink’s arrival in 2016 did not make the external cluster fabric optional; it created a boundary that every large system now has to design around explicitly rather than a replacement for what sat on the other side of it. The honest reading of four decades of this is not a relay race with one winner crossing each finish line, but an accumulation: every new fabric answers the specific, dated problem in front of it, and the ones it does not solve keep running alongside it, which is exactly why an AI training cluster commissioned today typically contains all three lineages described here at once, each doing the part of the job it was actually built for.
What the Ultra Ethernet Consortium’s specification does or does not displace by the end of this decade is not yet part of the historical record; only its founding and its first published specification are. Everything datable in this account up to that point shares one property worth stating plainly: each fabric is traceable to a named organization, a specific date, and a document or product that can still be checked against what its authors actually claimed for it at the time, rather than against what later marketing or later rivalry made of it in hindsight. That is a modest thing to be able to say about a technology history, and it is also the only kind of claim this account has tried to make.