GPU server rack with high-density compute nodes for AI training and inference workloads

Why TCO Matters More Than Hourly Price

The most common mistake in GPU hosting decisions is comparing sticker prices without accounting for the full cost picture. A cloud GPU instance at $3.00 per hour looks expensive compared to the $0.08/kWh electricity powering a colocated server. But the cloud price includes hardware amortization, maintenance, networking, cooling, and staff -- costs that are real but invisible when they appear in separate budget lines for colocation or on-premises deployments.

Total cost of ownership captures every dollar spent to deliver a GPU-hour of compute, regardless of where it appears in the accounting system. For organizations spending $100,000 or more per year on GPU infrastructure, a rigorous TCO analysis can reveal savings of 30 to 60 percent by choosing the right deployment model for their specific workload pattern.

This guide breaks down the complete cost structure for each deployment model -- cloud, colocation, and on-premises -- using current 2026 pricing for NVIDIA H100 and Blackwell-generation hardware.

The Three Deployment Models

Cloud GPU Hosting

Cloud providers (AWS, Google Cloud, Azure, CoreWeave, Lambda, and others) offer GPU instances billed by the hour, with reserved instance discounts for committed usage. The customer owns nothing physical. All hardware, power, cooling, networking, and facilities costs are bundled into the per-hour price.

Cloud pricing for NVIDIA H100 80GB SXM5 instances in late 2026 ranges from approximately $2.00 to $3.50 per GPU-hour on demand, and $1.20 to $2.00 per GPU-hour on 1-year reserved commitments. Blackwell B200 instances are priced at roughly 30 to 50 percent premiums over H100 for equivalent GPU-hours, reflecting both higher performance and constrained supply.

Colocation GPU Hosting

In colocation, the customer purchases and owns the GPU hardware and places it in a third-party data center that provides power, cooling, physical security, and network connectivity. The customer pays for rack space and power on a monthly basis, plus one-time or recurring costs for cross-connects, remote hands, and additional services.

Colocation costs for GPU-dense deployments in 2026 typically include:

  • Power: $0.08 to $0.15 per kWh (varies widely by region and contract)
  • Space: $150 to $500 per kW per month (including cooling), often quoted as a blended rate
  • Cross-connects: $200 to $500 per month per connection
  • Remote hands: $50 to $150 per incident or included in contract

On-Premises GPU Hosting

On-premises deployment places GPU hardware in the organization's own facility. This provides maximum control but requires the organization to build and operate the complete infrastructure stack: power distribution, cooling systems, physical security, fire suppression, and networking.

On-premises is the only option for organizations with strict data sovereignty requirements that preclude third-party hosting, or for workloads that generate or consume data volumes too large to move efficiently over external networks.

Full TCO Cost Breakdown

The following analysis uses a reference deployment of 64 NVIDIA H100 SXM5 GPUs (8 nodes of 8 GPUs each) operating at 80 percent average utilization over a 3-year period. All costs are in USD.

Capital Expenditures (CapEx)

Cost Category Cloud Colocation On-Premises
GPU servers (8x DGX H100 or equivalent) $0 (included in hourly rate) $1,600,000 - $2,400,000 $1,600,000 - $2,400,000
Networking (InfiniBand switches, cabling) $0 $80,000 - $150,000 $80,000 - $150,000
Facility infrastructure (power, cooling, racks) $0 $0 (included in colo fees) $300,000 - $800,000
Installation and commissioning $0 $15,000 - $30,000 $50,000 - $100,000
Total CapEx $0 $1,695,000 - $2,580,000 $2,030,000 - $3,450,000

Monthly Operating Expenditures (OpEx)

Cost Category Cloud Colocation On-Premises
Compute (GPU-hours at 80% utilization) $92,000 - $161,000/mo (on-demand) or $55,000 - $92,000/mo (reserved) $0 (hardware owned) $0 (hardware owned)
Power (64 GPUs + overhead at PUE 1.3) Included $5,200 - $9,700/mo $3,600 - $6,500/mo
Space and cooling Included $7,500 - $25,000/mo $2,000 - $5,000/mo (maintenance)
Networking / connectivity $500 - $3,000/mo (egress fees) $600 - $2,000/mo (cross-connects + bandwidth) $1,000 - $5,000/mo (ISP circuits)
Hardware maintenance / support Included $3,000 - $8,000/mo $3,000 - $8,000/mo
Staff (operations, engineering) $0 - $5,000/mo (DevOps only) $5,000 - $15,000/mo (partial FTE) $15,000 - $40,000/mo (facilities + IT staff)
Total Monthly OpEx $55,500 - $169,000 $21,300 - $59,700 $24,600 - $64,500

3-Year Total Cost of Ownership

Model 3-Year TCO (Low Estimate) 3-Year TCO (High Estimate) Cost per GPU-Hour
Cloud (reserved instances) $1,998,000 $3,312,000 $1.18 - $1.96
Colocation (owned hardware) $2,461,800 $4,729,200 $1.46 - $2.80
On-Premises $2,915,600 $5,772,000 $1.73 - $3.42

Important note: These ranges reflect significant variation in regional power costs, hardware procurement pricing, and contract terms. Your specific TCO will depend on your power rate, utilization pattern, hardware generation, and negotiated pricing. The key insight is not the specific numbers but the cost structure differences between models. At high utilization over 3 years, colocation typically delivers the lowest cost per GPU-hour. At variable or short-duration usage, cloud is more economical despite higher per-hour rates.

The Utilization Variable: When Each Model Wins

Utilization rate is the single most important variable in GPU hosting TCO. Hardware that you own costs the same whether it runs at 20 percent or 95 percent utilization. Cloud resources cost nothing when they are off.

This creates clear crossover points:

  • Below 40% utilization: Cloud wins almost always. You are paying only for what you use, and the hardware sits idle more than half the time. Buying hardware for this workload pattern wastes capital.
  • 40-60% utilization: The competitive zone. Cloud reserved instances and colocation are often comparable on a TCO basis. The choice depends on other factors: time to deploy, operational complexity tolerance, and whether you need the specific hardware configurations that colocation allows.
  • Above 60% utilization: Colocation begins to win clearly, and the advantage grows with utilization. At 80%+ sustained utilization over 2+ years, colocation is typically 40-60% less expensive than cloud.
  • Above 90% utilization: Both colocation and on-premises deliver significant savings over cloud. The choice between them depends on scale -- on-premises infrastructure makes financial sense when the deployment is large enough to amortize the facility costs across enough hardware.

Hidden Cost Factors That Change the Analysis

GPU Hardware Depreciation and Residual Value

GPU hardware purchased for colocation or on-premises deployment depreciates, but it retains residual value. A fleet of NVIDIA H100 servers purchased in 2024 still has significant resale or redeployment value in 2026. This residual value -- typically 20 to 40 percent of the original purchase price after 3 years -- reduces the effective CapEx when factored into the TCO calculation.

Cloud customers have no hardware residual value. Every dollar spent on cloud GPU compute is a pure operating expense with zero recovery at the end of the commitment period.

Time to Deploy

Cloud GPU instances can be provisioned in minutes (if capacity is available). Colocation deployment from hardware order to rack and power typically takes 4 to 12 weeks, including procurement lead times, shipping, installation, burn-in testing, and network configuration. On-premises deployment for a new facility can take 6 to 18 months including construction, permitting, and utility interconnection.

The cost of delayed deployment is real. If a cloud customer can start generating revenue from an AI product 3 months before a colocation deployment goes live, those 3 months of revenue can offset a significant portion of the cloud cost premium.

Power Cost Variability

Power is the largest ongoing cost in colocation and on-premises GPU hosting. A 64-GPU cluster consuming approximately 60 kW of IT load (plus cooling overhead) uses roughly 55,000 to 70,000 kWh per month. The difference between $0.07/kWh and $0.15/kWh is $4,400 to $5,600 per month -- a 10 to 15 percent swing in total colocation OpEx from power pricing alone.

This is why power-advantaged locations command significant interest for GPU infrastructure. The Rax Energy division focuses on securing low-cost power through power purchase agreements and renewable energy integration to minimize this cost component for hosted customers.

Networking and Data Transfer

Cloud egress fees are among the most criticized hidden costs in cloud computing. Moving large datasets out of cloud environments can cost $0.05 to $0.12 per GB. For AI workloads that generate large model outputs, training artifacts, or processed data, egress fees can add 5 to 15 percent to the effective cloud cost.

Colocation facilities offer flat-rate bandwidth pricing and free or low-cost cross-connect services between racks within the same facility. Data transfer between your own equipment in the same colocation hall is essentially free.

Operational Complexity and Staffing

Cloud offloads nearly all infrastructure operations to the provider. Colocation requires hardware procurement, deployment, monitoring, and maintenance -- but the facility handles power, cooling, and physical security. On-premises requires the full stack of infrastructure operations.

Staff costs are often underestimated. A small GPU colocation deployment might share an infrastructure engineer with other projects (partial FTE). A large on-premises deployment typically requires dedicated facilities staff for 24/7 coverage, which means a minimum of 4 to 5 full-time employees when factoring in shifts, vacations, and training.

Decision Framework: Choosing the Right Model

Rather than a one-size-fits-all recommendation, the right deployment model depends on your specific situation across several dimensions:

Factor Favors Cloud Favors Colocation Favors On-Premises
Utilization pattern Variable, below 50% Steady, above 60% Near-continuous, above 85%
Deployment horizon Under 18 months 2 to 5 years 5+ years
Scale (GPU count) 1 to 32 GPUs 8 to 1,000+ GPUs 500+ GPUs (to amortize facility)
Time to deploy Minutes to hours 4 to 12 weeks 6 to 18 months
Data sovereignty Provider regions Chosen facility Full control
Operational complexity Minimal (provider managed) Moderate (hardware ops) High (full stack)
Capital availability OpEx-preferred CapEx available Large CapEx available

The Hybrid Approach

Many organizations are adopting hybrid strategies that combine two or all three deployment models:

  • Colocation base + cloud burst: Own enough hardware in colocation to handle the steady-state workload, and burst to cloud for peak demand periods. This captures the cost efficiency of owned hardware at high utilization while retaining the flexibility to scale temporarily without over-provisioning.
  • Cloud for experimentation + colocation for production: Use cloud instances for early-stage model development, experimentation, and evaluation where utilization is sporadic. Once a model reaches production and requires sustained GPU allocation, migrate to colocated infrastructure for long-term cost optimization.
  • Multi-cloud with colocation anchor: Maintain a colocation presence as the cost-efficient anchor, use multiple cloud providers for geographic distribution and redundancy, and leverage Kubernetes GPU orchestration to shift workloads dynamically based on cost, availability, and latency requirements.

Key takeaway: There is no universally "cheapest" GPU hosting option. The optimal model depends on your utilization pattern, deployment horizon, operational capabilities, and capital structure. For organizations running sustained GPU workloads over multi-year periods, colocation typically delivers the lowest total cost of ownership per GPU-hour -- often 40 to 60 percent below cloud pricing at equivalent scale and utilization.

Frequently Asked Questions

How much does GPU hosting cost per month in colocation vs cloud?

Cloud GPU hosting for a single NVIDIA H100 instance typically costs $1,800 to $2,900 per month on a reserved commitment. Colocation hosting for equivalent hardware you own costs approximately $800 to $1,500 per month in power, space, and connectivity fees, plus the upfront hardware purchase of $25,000 to $40,000 per GPU. Over 3 years at high utilization, colocation is typically 40 to 60 percent less expensive on a TCO basis.

When does cloud GPU hosting make more financial sense than colocation?

Cloud is more cost-effective when utilization is below 40 to 50 percent, workloads are highly variable or burst-oriented, the deployment horizon is under 12 to 18 months, or the organization lacks staff to manage hardware. Cloud also wins when access to the latest GPU generation is critical and hardware lead times exceed 6 months.

What are the hidden costs of on-premises GPU hosting?

Often underestimated costs include electrical infrastructure upgrades (transformers, switchgear, PDUs for high-density racks), cooling system modifications, structural reinforcement for heavy GPU servers, dedicated InfiniBand networking equipment, 24/7 facilities staff, and the opportunity cost of building space used for IT rather than revenue-generating activities.

How does power cost affect GPU hosting TCO?

Power typically accounts for 30 to 50 percent of ongoing operational expenses. A single H100 consumes approximately 500 kWh per month at full load. At $0.10/kWh, that is $50 per month per GPU in direct power alone, plus $15 to $25 for cooling overhead. The difference between a $0.07/kWh and $0.15/kWh facility rate can swing total colocation OpEx by 10 to 15 percent.

What is the break-even point between cloud and colocation GPU hosting?

For a cluster of 8 H100 GPUs at 80 percent utilization, the break-even between cloud on-demand and colocation with owned hardware occurs at approximately 14 to 18 months. With cloud reserved instances, the break-even extends to 20 to 26 months. Organizations planning 3+ year deployments with steady utilization will find colocation TCO 40 to 60 percent lower than cloud.

Evaluate Your GPU Hosting Options

Rax Data & Energy provides GPU colocation and AI compute hosting with competitive power rates, enterprise-grade redundancy, and the infrastructure density required for modern AI workloads.

Get a Quote View Pricing

Related Articles