Cluster fabrics: Ethernet, InfiniBand and the fight for the back end
Beyond the rack, accelerators talk over a switched network spanning a data hall. The requirement is unusual: not high average throughput but consistently low tail latency and no dropped packets, because one straggler stalls every node in a synchronised training step.
In one sentence
A cluster fabric is the switched network connecting accelerator nodes across a data centre, engineered for lossless, low-tail-latency collective communication rather than for general internet traffic.
Ordinary data-centre networks are built on the assumption that occasional congestion and retransmission are fine. Training traffic breaks that assumption. Every step ends in a collective operation in which all nodes must exchange results, so a single slow or retransmitting link sets the pace for the entire cluster.
Two families answer this. A specialised high-performance interconnect was built lossless from the start and dominated early large clusters. Ethernet, with congestion-control and scheduling extensions, has been catching up — helped by an industry consortium and by buyers who prefer a technology with many vendors and an existing operational skill base.
How it works
What 'lossless' actually requires
Flow control that tells a sender to pause before a buffer overflows, congestion signalling that reacts before queues build, and spreading traffic across many equal-cost paths at fine granularity rather than pinning a flow to one path. Get these wrong and the network is fast on paper and slow in a collective operation.
Front-end and back-end networks
AI clusters run two distinct networks. The front end carries storage, management and general traffic and looks like a normal data-centre network. The back end carries only accelerator-to-accelerator traffic, has far more bandwidth per node, and is where the specialised engineering goes. They are separately designed and separately budgeted.
Why the topology is chosen carefully
The fabric's shape decides how many switch hops and optical links a message crosses. Since optics dominate both cost and power in the network, a topology that reduces hops is a direct cost saving — which is why rail-optimised layouts and other AI-specific arrangements have replaced the generic fat tree in these builds.
What this depends on
2 of these are marked as a chokepoint: a handful of qualified suppliers, a multi-year lead time, or a single geography.
Supply chainChokepoint
Merchant switch silicon
Almost every switch in an AI cluster is built around a small number of high-radix switch chips from a small number of suppliers.
Ethernet only behaves acceptably for this traffic because of an agreed set of extensions; whether vendors implement them compatibly is the practical question.
Technology
Collective communication libraries
The software that decides how a collective operation is decomposed and scheduled determines how much of the network's capability is realised.
Links inside a rack and to the nearest switch are copper wherever reach allows, because each one avoided is a pair of transceivers not bought and not powered.
Tens of thousands of fibre pairs have to be pulled, terminated, tested and documented before a cluster runs. That work is on the building's critical path and cannot be compressed much by spending more.
What depends on this
Other pages in this map that name Cluster fabrics as something they cannot do without.
What each company supplies at this step, and — where a public figure exists — its share of this specific market — with what that share measures, the period it covers and who published it. Some rows also show the company’s own reported revenue for the segment covering this step, which is a different thing: it says how much this business matters to that company, not how much of the market it holds. Not a ranking and not a recommendation.
Runs and publishes the largest optically switched fabric, which is where much of the field's design vocabulary came from.
Cornelis NetworksPrivate
Sells a fabric engineered specifically for collectives, the surviving alternative to Ethernet and InfiniBand.
What would change the picture
Whether Ethernet's share of new back-end fabrics keeps rising as the standardised extensions mature.
Whether network cost and power stay near their current share of a cluster as per-node bandwidth grows.
Whether topology innovations meaningfully reduce the optics count per accelerator.
Questions people ask about this
Why can't a normal data-centre network be used?
Because of what training traffic does. It is bursty, synchronised and all-to-all, and every step waits for the slowest participant. A network that handles average load well but occasionally drops or delays packets converts into idle accelerators, which is the most expensive kind of idle there is.
Is the network really a large share of cluster cost?
It is a material one — commonly quoted in the range of a tenth of the capital cost of an AI cluster, most of it in optics rather than switches. It is also a real share of power, which is why co-packaged optics and topologies that reduce link count get so much attention.
Each page explains one technology in plain language, states what it depends on, and names companies by what they supply at that step. Company roles are described qualitatively and deliberately carry no market shares, revenue figures or rankings — those change faster than an explainer can, and a stale number is worse than none. Ticker links point at company pages on this site and are provided for reference only.
Nothing here is investment advice, a recommendation, or a forecast. A company named on a page about a technology is not thereby a good investment, and the chokepoints described are structural facts about supply chains rather than predictions about prices. Technology moves; where a page describes something as unresolved or in development, that was true when it was written.