Data Centers

HPC Colocation: High-Performance Computing Infrastructure for Data Centers

High-performance computing data center with GPU server racks and liquid cooling infrastructure

High-performance computing (HPC) workloads demand infrastructure that traditional enterprise data centers simply cannot provide. From computational fluid dynamics simulations to genomics sequencing, weather modeling, and large-scale AI training, HPC applications require extreme compute density, ultra-low-latency interconnects, and specialized cooling that pushes far beyond conventional IT hosting.

HPC colocation offers a compelling alternative to building private facilities. By placing compute hardware in purpose-built data centers, organizations gain access to high-density power, advanced cooling, and carrier-neutral connectivity without the capital expenditure and 18-24 month timeline of ground-up construction.

What Makes HPC Colocation Different

Standard colocation facilities are designed for 5-10 kW per rack. HPC workloads routinely require 30 to 100+ kW per rack, creating fundamentally different demands across every infrastructure layer:

Compute Infrastructure Requirements

GPU Clusters

Modern HPC increasingly relies on GPU-accelerated computing. A single NVIDIA GB200 NVL72 rack can draw 120 kW, requiring liquid cooling and high-density power infrastructure that few facilities can provide.

Key GPU cluster considerations for colocation:

CPU Nodes

Many HPC workloads remain CPU-bound, particularly in computational chemistry, finite element analysis, and weather modeling. Modern HPC CPU nodes use AMD EPYC 9004 (Genoa/Bergamo) or Intel Xeon Sapphire Rapids processors with high core counts (up to 128 cores per socket) and large L3 caches optimized for memory-intensive parallel computation.

Network Fabric Design

The network fabric is the single most critical infrastructure decision for HPC colocation. Tightly coupled parallel workloads spend significant time in inter-node communication, and network latency directly determines job completion time.

InfiniBand

InfiniBand remains the gold standard for HPC interconnects:

Ultra Ethernet

The Ultra Ethernet Consortium (UEC) is developing Ethernet-based alternatives targeting HPC and AI workloads. While not yet matching InfiniBand's latency, RoCEv2 (RDMA over Converged Ethernet) at 400 GbE provides a cost-effective option for loosely coupled workloads.

Storage Tiers for HPC

HPC workloads generate and consume massive datasets. A well-designed storage architecture uses multiple tiers:

Storage networking must match compute fabric performance. NVMe-oF (NVMe over Fabrics) enables remote NVMe access at near-local latency, critical for data-intensive HPC applications.

Power and Cooling for High-Density HPC

Power and cooling represent the most significant infrastructure challenge and cost driver for HPC colocation.

Power Architecture

HPC facilities require robust power redundancy with:

Cooling Technologies

At 30+ kW per rack, air cooling alone is insufficient. Modern HPC colocation facilities deploy:

Choosing an HPC Colocation Provider

Evaluating colocation providers for HPC requires a different checklist than standard enterprise hosting. Use this colocation buyer's checklist as a starting point, then add HPC-specific criteria:

  1. Power density commitment: Can the facility deliver 50+ kW per rack today, with a path to 100+ kW? Verify with actual deployed customer references, not just marketing specifications.
  2. Cooling proof: Ask for PUE data at HPC densities (not blended facility average). Target PUE below 1.2 for liquid-cooled deployments.
  3. Cross-connect capabilities: Can you deploy InfiniBand between your cabinets? Some facilities restrict non-Ethernet cabling.
  4. Floor loading and structural capacity: GPU-dense racks weigh 2,000-3,000+ lbs. Verify floor ratings and reinforcement.
  5. Hands-and-eyes support: HPC hardware requires specialized remote hands. GPU re-seating, liquid cooling loop maintenance, and InfiniBand cable management differ from standard IT operations.
  6. Contract flexibility: HPC projects often have defined timelines. Look for term flexibility rather than rigid 3-5 year commitments.

Cost Considerations

HPC colocation pricing differs significantly from standard colocation. Key cost components:

Cost Component Standard Colo HPC Colo
Power (per kW/month) $100-150 $200-500+
Cooling premium Included in PUE 15-30% surcharge for liquid
Cross-connects $200-500/mo each $500-2,000/mo (InfiniBand)
Remote hands $75-150/hour $150-300/hour (specialized)

Despite higher per-unit costs, total cost of ownership (TCO) for HPC colocation is typically 30-50% lower than building a private facility when accounting for construction costs, staffing, and the time value of faster deployment (weeks vs. 18-24 months).

Security and Compliance

Many HPC workloads involve sensitive data: government research, pharmaceutical R&D, financial modeling, and defense applications. Verify that your colocation provider offers:

Future Trends in HPC Colocation

The HPC colocation market is evolving rapidly:

Need HPC Colocation?

Rax Data & Energy offers high-density colocation infrastructure with liquid cooling, InfiniBand support, and flexible power commitments for HPC workloads.

Get a Quote

Frequently Asked Questions

What is HPC colocation?

HPC colocation is a hosting model where organizations place high-performance computing hardware, including GPU clusters, CPU nodes, and high-speed interconnects, in a third-party data center that provides power, cooling, physical security, and network connectivity optimized for compute-intensive workloads.

What power density do HPC workloads require?

Modern HPC workloads typically require 30 to 100+ kW per rack, far exceeding the 5-10 kW per rack of standard enterprise IT. GPU-dense configurations with NVIDIA H100 or H200 GPUs can reach 70-120 kW per rack, requiring liquid cooling and high-density power distribution.

What networking does HPC colocation need?

HPC workloads require ultra-low-latency interconnects like InfiniBand NDR (400 Gbps) or NVIDIA Quantum switches for GPU-to-GPU communication. Standard Ethernet is insufficient for tightly coupled parallel workloads. The network fabric design directly impacts job completion time.

How much does HPC colocation cost?

HPC colocation costs vary widely based on power density and cooling requirements. Expect $200 to $500+ per kW per month for high-density hosting, plus interconnect and storage charges. Total cost of ownership is typically 30-50% lower than building a private facility for most organizations.

What cooling is needed for HPC colocation?

HPC systems generating 30+ kW per rack typically require direct liquid cooling (DLC), rear-door heat exchangers, or immersion cooling. Traditional raised-floor air cooling cannot handle the thermal loads of modern GPU clusters. Many facilities use a hybrid approach combining liquid cooling for compute nodes with air cooling for storage and networking.