How AI Datacenter Interconnects Actually Work
Underneath every all-reduce sits a queue pair, a doorbell register, a PAM4 laser and a switch buffer that is either draining or filling. This is the mechanism, one layer down from the equipment.
Tagged collective communication · show all articles
Underneath every all-reduce sits a queue pair, a doorbell register, a PAM4 laser and a switch buffer that is either draining or filling. This is the mechanism, one layer down from the equipment.
A benchmark result is reached once, under conditions someone else chose. Production throughput means choosing precision on purpose, profiling the real bottleneck, tuning collectives, and planning capacity for a fleet that fails the ordinary way.
Scale-up fabrics move data at the speed of a shared memory space; scale-out fabrics move it at the speed of a switch queue. A training job lives across both, and where the boundary falls decides how fast the cluster actually trains.
Once a model no longer fits on one device, the interconnect stops being plumbing and becomes the design surface. Every way of splitting a model generates its own traffic, and the fabric decides which splits are affordable.