GPU and AI infrastructure built around the workload.
GPU systems are specified from the workload inward. Model size, batch behaviour, dataset throughput and rack power decide the platform long before a part number does.
Not sure what to specify?
Start with the workload and operating constraints. The accelerator, system design, network, and storage can follow from those needs.
Start from the workload
Training, fine-tuning and inference place different demands on the same hardware. Training rewards accelerator interconnect bandwidth and sustained storage throughput. Inference rewards memory capacity per accelerator, latency and predictable density. HPC workloads are often bound by the fabric rather than by the accelerator.
- LLM training and fine-tuning clusters
- LLM and general model inference infrastructure
- Computer vision and signal processing workloads
- Scientific and engineering HPC with GPU acceleration
- Mixed development and experimentation environments
Requirements that decide the platform
| Requirement | What to state |
|---|---|
| Accelerator count and memory | Per-system count and the memory capacity each workload needs. |
| Accelerator platform | Whether PCIe cards or a baseboard/module platform is expected or open. |
| Interconnect | Whether accelerators must communicate over a high-bandwidth link. |
| Cluster fabric | Ethernet or InfiniBand, port speed, and how many nodes will scale. |
| Storage throughput | Sustained read bandwidth the training pipeline has to sustain. |
| Rack power and cooling | kW per rack available, and whether liquid cooling is possible. |
Power and facility constraints come first
Dense GPU systems frequently exceed the power and cooling envelope of the rack they were intended for. Stating available kW per rack, inlet temperature and whether liquid cooling is an option prevents quotations for systems the site cannot host.
If the facility is the constraint, a lower-density configuration across more racks is often the faster path to a working deployment.
Storage and networking are part of the AI requirement
An accelerator fleet that starves on data is an expensive idle asset. Specify the storage tier feeding the pipeline and the fabric between nodes alongside the compute, not after it.
How 46 Systems handles AI infrastructure requirements
Describe the workload and the constraints, and 46 Systems organizes sourcing across manufacturers and platform types rather than defaulting to one. Where accelerator supply is constrained, alternatives and staged delivery options can be requested explicitly.
Things worth considering
- Choosing an accelerator before defining model size and throughput targets
- Ignoring rack power until after the systems are quoted
- Sizing GPUs without sizing the data pipeline behind them
- Assuming Ethernet and InfiniBand are interchangeable at scale
- Planning no headroom for the next model or dataset
Related pages
Describe your AI infrastructure requirement.
Tell us the workload and constraints, or configure a GPU server directly.