GPU server buying guide.
GPU systems are bought on accelerator specifications and fail on everything around them — power, cooling, host memory and the storage feeding the job.
Size the accelerator memory first
Accelerator memory decides what will run at all. Model weights, activations, optimizer state and batch size all live there, and exceeding it turns a working job into an out-of-memory error rather than a slow one.
For inference, memory capacity per accelerator usually matters more than raw throughput. For training, both matter, and interconnect bandwidth between accelerators becomes a third constraint.
PCIe cards or a module platform
| Platform | Fits | Watch for |
|---|---|---|
| PCIe accelerator cards | Flexible node counts, mixed workloads, smaller clusters | Slot budget, per-card power, airflow across cards |
| Baseboard / module platforms | Dense training, tightly coupled multi-accelerator jobs | Rack power, cooling method, chassis availability |
Do not starve the accelerators
- Host CPU cores to keep data loaders ahead of the accelerators
- Host memory typically well above total accelerator memory
- Local NVMe scratch sized for the working dataset
- Network bandwidth matched to the shared dataset tier
- Storage read throughput measured as sustained, not peak
Power and cooling are the real limit
A dense GPU node can draw more than a whole legacy rack was provisioned for. Before selecting a platform, confirm kW available per rack, the inlet temperature, and whether liquid cooling is possible at the site. If it is not, plan lower density across more racks rather than discovering the constraint after delivery.
Plan for the next model
Accelerator requirements change faster than hardware refresh cycles. Leaving headroom in rack power, fabric port count and storage capacity costs far less at purchase than retrofitting a full rack later.
Describe your GPU requirement.
Tell us the workload and constraints, or configure a GPU server.