High-performance computing (HPC) workloads demand infrastructure that traditional enterprise data centers simply cannot provide. From computational fluid dynamics simulations to genomics sequencing, weather modeling, and large-scale AI training, HPC applications require extreme compute density, ultra-low-latency interconnects, and specialized cooling that pushes far beyond conventional IT hosting.
HPC colocation offers a compelling alternative to building private facilities. By placing compute hardware in purpose-built data centers, organizations gain access to high-density power, advanced cooling, and carrier-neutral connectivity without the capital expenditure and 18-24 month timeline of ground-up construction.
What Makes HPC Colocation Different
Standard colocation facilities are designed for 5-10 kW per rack. HPC workloads routinely require 30 to 100+ kW per rack, creating fundamentally different demands across every infrastructure layer:
- Power distribution: High-amperage circuits (60A-100A per cabinet), redundant PDUs rated for sustained high loads, and busway distribution replacing traditional whip-and-plug architectures
- Cooling: Direct liquid cooling (DLC), rear-door heat exchangers, or immersion cooling rather than raised-floor air handling
- Structural: Floor loading capacity of 250+ lbs/sq ft versus the standard 150 lbs/sq ft, reinforced cable trays, and wider hot/cold aisles
- Networking: InfiniBand or RDMA-capable Ethernet fabrics with sub-microsecond latency, not just standard 10/25 GbE
- Fire suppression: Clean agent systems compatible with liquid cooling loops and high-value GPU hardware
Compute Infrastructure Requirements
GPU Clusters
Modern HPC increasingly relies on GPU-accelerated computing. A single NVIDIA GB200 NVL72 rack can draw 120 kW, requiring liquid cooling and high-density power infrastructure that few facilities can provide.
Key GPU cluster considerations for colocation:
- NVIDIA H100/H200 SXM configurations draw 40-70 kW per rack depending on density
- GB200 NVL72 racks require direct liquid cooling loops with 45-50°C supply water temperatures
- GPU memory bandwidth (HBM3/HBM3e) determines throughput for LLM training and inference
- NVLink and NVSwitch interconnects provide 900 GB/s GPU-to-GPU bandwidth within a node
CPU Nodes
Many HPC workloads remain CPU-bound, particularly in computational chemistry, finite element analysis, and weather modeling. Modern HPC CPU nodes use AMD EPYC 9004 (Genoa/Bergamo) or Intel Xeon Sapphire Rapids processors with high core counts (up to 128 cores per socket) and large L3 caches optimized for memory-intensive parallel computation.
Network Fabric Design
The network fabric is the single most critical infrastructure decision for HPC colocation. Tightly coupled parallel workloads spend significant time in inter-node communication, and network latency directly determines job completion time.
InfiniBand
InfiniBand remains the gold standard for HPC interconnects:
- NDR (400 Gbps): Current generation, sub-microsecond latency, RDMA native
- XDR (800 Gbps): Next generation arriving in late 2026, doubling bandwidth
- Fat-tree or Dragonfly topologies minimize hop count and congestion
- Adaptive routing and congestion management via NVIDIA Quantum switches
Ultra Ethernet
The Ultra Ethernet Consortium (UEC) is developing Ethernet-based alternatives targeting HPC and AI workloads. While not yet matching InfiniBand's latency, RoCEv2 (RDMA over Converged Ethernet) at 400 GbE provides a cost-effective option for loosely coupled workloads.
Storage Tiers for HPC
HPC workloads generate and consume massive datasets. A well-designed storage architecture uses multiple tiers:
- Burst buffer (NVMe tier): All-flash NVMe arrays for checkpoint/restart and scratch space. 100+ GB/s aggregate throughput.
- Parallel file system: Lustre, GPFS/Spectrum Scale, or BeeGFS for shared project storage. Scales to petabytes with 50+ GB/s sustained throughput.
- Archive tier: Object storage (Ceph, MinIO) or tape for long-term dataset retention at lower cost per terabyte.
Storage networking must match compute fabric performance. NVMe-oF (NVMe over Fabrics) enables remote NVMe access at near-local latency, critical for data-intensive HPC applications.
Power and Cooling for High-Density HPC
Power and cooling represent the most significant infrastructure challenge and cost driver for HPC colocation.
Power Architecture
HPC facilities require robust power redundancy with:
- N+1 or 2N UPS configurations sized for sustained high loads (not just brief peaks)
- High-efficiency transformers and switchgear rated for continuous operation at 80%+ load
- Intelligent PDUs with per-outlet monitoring and remote management
- Generator capacity with block loading to handle instant full-site failover
Cooling Technologies
At 30+ kW per rack, air cooling alone is insufficient. Modern HPC colocation facilities deploy:
- Direct-to-chip liquid cooling (DLC): Warm water (35-45°C) circulated directly to CPU/GPU cold plates. Removes 70-80% of heat at the source.
- Immersion cooling: Single-phase or two-phase dielectric fluid submerges entire servers. Enables extreme density but requires specialized maintenance procedures.
- Rear-door heat exchangers: Water-cooled rear doors capture exhaust heat. Effective for 20-40 kW per rack without modifying server internals.
- Hybrid approaches: Liquid cooling for compute nodes, air cooling for storage and networking equipment.
Choosing an HPC Colocation Provider
Evaluating colocation providers for HPC requires a different checklist than standard enterprise hosting. Use this colocation buyer's checklist as a starting point, then add HPC-specific criteria:
- Power density commitment: Can the facility deliver 50+ kW per rack today, with a path to 100+ kW? Verify with actual deployed customer references, not just marketing specifications.
- Cooling proof: Ask for PUE data at HPC densities (not blended facility average). Target PUE below 1.2 for liquid-cooled deployments.
- Cross-connect capabilities: Can you deploy InfiniBand between your cabinets? Some facilities restrict non-Ethernet cabling.
- Floor loading and structural capacity: GPU-dense racks weigh 2,000-3,000+ lbs. Verify floor ratings and reinforcement.
- Hands-and-eyes support: HPC hardware requires specialized remote hands. GPU re-seating, liquid cooling loop maintenance, and InfiniBand cable management differ from standard IT operations.
- Contract flexibility: HPC projects often have defined timelines. Look for term flexibility rather than rigid 3-5 year commitments.
Cost Considerations
HPC colocation pricing differs significantly from standard colocation. Key cost components:
| Cost Component | Standard Colo | HPC Colo |
|---|---|---|
| Power (per kW/month) | $100-150 | $200-500+ |
| Cooling premium | Included in PUE | 15-30% surcharge for liquid |
| Cross-connects | $200-500/mo each | $500-2,000/mo (InfiniBand) |
| Remote hands | $75-150/hour | $150-300/hour (specialized) |
Despite higher per-unit costs, total cost of ownership (TCO) for HPC colocation is typically 30-50% lower than building a private facility when accounting for construction costs, staffing, and the time value of faster deployment (weeks vs. 18-24 months).
Security and Compliance
Many HPC workloads involve sensitive data: government research, pharmaceutical R&D, financial modeling, and defense applications. Verify that your colocation provider offers:
- Multi-layer physical security: biometric access, mantrap entries, 24/7 security staff, CCTV with 90-day retention
- Compliance certifications relevant to your workload: SOC 2 Type II, ISO 27001, FedRAMP (for government), HIPAA (for healthcare)
- Dedicated cages or private suites for hardware isolation
- Network segmentation and private interconnects between your cabinets
Future Trends in HPC Colocation
The HPC colocation market is evolving rapidly:
- Convergence with AI infrastructure: The line between HPC and AI workloads is blurring. Facilities designed for one increasingly support both.
- Sustainable cooling: Waste heat reuse, free cooling in favorable climates, and efficiency improvements are reducing environmental impact.
- Edge HPC: Edge data centers are beginning to offer HPC-class infrastructure for latency-sensitive applications.
- As-a-service models: GPUaaS and HPC-as-a-service offerings reduce capital requirements while maintaining bare-metal performance.
- Quantum-HPC hybrid: Early quantum computing systems are being deployed alongside classical HPC clusters, requiring specialized environmental controls.
Need HPC Colocation?
Rax Data & Energy offers high-density colocation infrastructure with liquid cooling, InfiniBand support, and flexible power commitments for HPC workloads.
Get a QuoteFrequently Asked Questions
What is HPC colocation?
HPC colocation is a hosting model where organizations place high-performance computing hardware, including GPU clusters, CPU nodes, and high-speed interconnects, in a third-party data center that provides power, cooling, physical security, and network connectivity optimized for compute-intensive workloads.
What power density do HPC workloads require?
Modern HPC workloads typically require 30 to 100+ kW per rack, far exceeding the 5-10 kW per rack of standard enterprise IT. GPU-dense configurations with NVIDIA H100 or H200 GPUs can reach 70-120 kW per rack, requiring liquid cooling and high-density power distribution.
What networking does HPC colocation need?
HPC workloads require ultra-low-latency interconnects like InfiniBand NDR (400 Gbps) or NVIDIA Quantum switches for GPU-to-GPU communication. Standard Ethernet is insufficient for tightly coupled parallel workloads. The network fabric design directly impacts job completion time.
How much does HPC colocation cost?
HPC colocation costs vary widely based on power density and cooling requirements. Expect $200 to $500+ per kW per month for high-density hosting, plus interconnect and storage charges. Total cost of ownership is typically 30-50% lower than building a private facility for most organizations.
What cooling is needed for HPC colocation?
HPC systems generating 30+ kW per rack typically require direct liquid cooling (DLC), rear-door heat exchangers, or immersion cooling. Traditional raised-floor air cooling cannot handle the thermal loads of modern GPU clusters. Many facilities use a hybrid approach combining liquid cooling for compute nodes with air cooling for storage and networking.