Plutux

Sub-system

Cluster fabrics: Ethernet, InfiniBand and the fight for the back end

Beyond the rack, accelerators talk over a switched network spanning a data hall. The requirement is unusual: not high average throughput but consistently low tail latency and no dropped packets, because one straggler stalls every node in a synchronised training step.

In one sentence

A cluster fabric is the switched network connecting accelerator nodes across a data centre, engineered for lossless, low-tail-latency collective communication rather than for general internet traffic.

Ordinary data-centre networks are built on the assumption that occasional congestion and retransmission are fine. Training traffic breaks that assumption. Every step ends in a collective operation in which all nodes must exchange results, so a single slow or retransmitting link sets the pace for the entire cluster.

Two families answer this. A specialised high-performance interconnect was built lossless from the start and dominated early large clusters. Ethernet, with congestion-control and scheduling extensions, has been catching up — helped by an industry consortium and by buyers who prefer a technology with many vendors and an existing operational skill base.

How it works

What 'lossless' actually requires

Flow control that tells a sender to pause before a buffer overflows, congestion signalling that reacts before queues build, and spreading traffic across many equal-cost paths at fine granularity rather than pinning a flow to one path. Get these wrong and the network is fast on paper and slow in a collective operation.

Front-end and back-end networks

AI clusters run two distinct networks. The front end carries storage, management and general traffic and looks like a normal data-centre network. The back end carries only accelerator-to-accelerator traffic, has far more bandwidth per node, and is where the specialised engineering goes. They are separately designed and separately budgeted.

Why the topology is chosen carefully

The fabric's shape decides how many switch hops and optical links a message crosses. Since optics dominate both cost and power in the network, a topology that reduces hops is a direct cost saving — which is why rail-optimised layouts and other AI-specific arrangements have replaced the generic fat tree in these builds.

What this depends on

2 of these are marked as a chokepoint: a handful of qualified suppliers, a multi-year lead time, or a single geography.

  • Supply chainChokepoint

    Merchant switch silicon

    Almost every switch in an AI cluster is built around a small number of high-radix switch chips from a small number of suppliers.

    Switch silicon
  • Supply chainChokepoint

    Optical transceivers

    A large cluster needs optical modules in the hundreds of thousands; they are a major share of network cost and power.

    Optical interconnect
  • Standard

    Congestion control and RDMA standards

    Ethernet only behaves acceptably for this traffic because of an agreed set of extensions; whether vendors implement them compatibly is the practical question.

  • Technology

    Collective communication libraries

    The software that decides how a collective operation is decomposed and scheduled determines how much of the network's capability is realised.

    Training frameworks
  • Supply chain

    Copper cabling for short hops

    Links inside a rack and to the nearest switch are copper wherever reach allows, because each one avoided is a pair of transceivers not bought and not powered.

    Cables and connectors
  • Resource

    Structured fibre plant and installation labour

    Tens of thousands of fibre pairs have to be pulled, terminated, tested and documented before a cluster runs. That work is on the building's critical path and cannot be compressed much by spending more.

What depends on this

Other pages in this map that name Cluster fabrics as something they cannot do without.

Who supplies this

What each company supplies at this step, and — where a public figure exists — its share of this specific market — with what that share measures, the period it covers and who published it. Some rows also show the company’s own reported revenue for the segment covering this step, which is a different thing: it says how much this business matters to that company, not how much of the market it holds. Not a ranking and not a recommendation.

  • Arista NetworksANET

    Supplies the high-radix Ethernet switches used in AI back-end fabrics.

  • BroadcomAVGO

    Supplies the merchant switch silicon most of those switches are built around.

  • NVIDIANVDA

    Supplies both the specialised high-performance interconnect and its own Ethernet platform for AI clusters.

  • Cisco SystemsCSCO

    Supplies switching and networking systems for AI data-centre build-outs.

  • CelesticaCLS

    Builds the switches themselves for hyperscalers that design their own.

  • Hewlett Packard EnterpriseHPE

    Supplies Juniper and Slingshot fabrics, the third stack alongside Ethernet and InfiniBand.

  • Marvell TechnologyMRVL

    Supplies switch and interface silicon to operators building their own fabric hardware.

  • Huawei TechnologiesPrivate

    Supplies the switches Chinese clusters are wired with, end to end and without merchant silicon.

  • Accton TechnologyTaiwan

    Manufactures the white-box switches operators buy when they do not want a brand vendor's software attached.

  • NokiaNOK

    Sells data-centre routing built on its own silicon into AI fabrics, and the wide-area links that join sites.

  • AlphabetGOOGL

    Runs and publishes the largest optically switched fabric, which is where much of the field's design vocabulary came from.

  • Cornelis NetworksPrivate

    Sells a fabric engineered specifically for collectives, the surviving alternative to Ethernet and InfiniBand.

What would change the picture

  • Whether Ethernet's share of new back-end fabrics keeps rising as the standardised extensions mature.

  • Whether network cost and power stay near their current share of a cluster as per-node bandwidth grows.

  • Whether topology innovations meaningfully reduce the optics count per accelerator.

Questions people ask about this

Why can't a normal data-centre network be used?
Because of what training traffic does. It is bursty, synchronised and all-to-all, and every step waits for the slowest participant. A network that handles average load well but occasionally drops or delays packets converts into idle accelerators, which is the most expensive kind of idle there is.
Is the network really a large share of cluster cost?
It is a material one — commonly quoted in the range of a tenth of the capital cost of an AI cluster, most of it in optics rather than switches. It is also a real share of power, which is why co-packaged optics and topologies that reduce link count get so much attention.

How these pages are written

Each page explains one technology in plain language, states what it depends on, and names companies by what they supply at that step. Company roles are described qualitatively and deliberately carry no market shares, revenue figures or rankings — those change faster than an explainer can, and a stale number is worse than none. Ticker links point at company pages on this site and are provided for reference only.

Nothing here is investment advice, a recommendation, or a forecast. A company named on a page about a technology is not thereby a good investment, and the chokepoints described are structural facts about supply chains rather than predictions about prices. Technology moves; where a page describes something as unresolved or in development, that was true when it was written.

Plutux는 투자자문업자가 아닙니다. 시장 데이터와 AI가 생성한 분석은 정보 제공 및 교육 목적일 뿐 투자 자문이 아닙니다. 면책 조항

© Plutux Technology Limited 2026