Ethernet vs InfiniBand for AI clusters.

The fabric decision is usually made on cluster scale and operational reality, not on a benchmark. Both options build working clusters; the difference is where the effort goes.

What actually differs

DimensionEthernetInfiniBand
EcosystemVery broad, many vendorsNarrower, specialist
OperationsFamiliar to most network teamsDistinct tooling and skills
Collective performanceStrong with careful lossless designPurpose-built for it
SupplyGenerally broader availabilityCan be constrained with GPU demand
ReuseIntegrates with existing infrastructureTypically a dedicated fabric

Where scale changes the answer

For a handful of nodes, either fabric works and the decision is mostly operational. As a training cluster grows and jobs depend on tightly synchronized collective operations across many nodes, the tail latency of the fabric starts to dominate job time, and that is where InfiniBand has traditionally been chosen.

Modern high-speed Ethernet with congestion control and careful design is used at large scale too. Both are legitimate; what is not legitimate is assuming they are interchangeable at the point of purchase.

Inference and mixed clusters

Inference fleets rarely need tightly coupled collective bandwidth between nodes. They need predictable north-south capacity to serve requests and enough bandwidth to load models. Ethernet is usually the straightforward answer there.

Practical specification checklist

  • Node count today, and the count the fabric must reach
  • Whether jobs span nodes or run within a single node
  • Existing switching and the team's operational familiarity
  • Port speed per node and total uplink capacity required
  • Whether optics and cabling are in scope for the quotation
  • Whether the fabric choice is fixed or open to recommendation

Describe your cluster requirement.

Compute and fabric, sourced as one requirement.