The single biggest bottleneck for AI workloads in modern data centers is not compute power. It is memory. Large language models with hundreds of billions of parameters require tens to hundreds of gigabytes of memory per GPU just to hold model weights, activations, and key-value caches during inference. Training runs demand even more. When a GPU's local high-bandwidth memory (HBM) is exhausted, operators face costly choices: split the model across more GPUs (adding inter-GPU communication overhead), reduce batch sizes (sacrificing throughput), or offload to slower storage tiers (destroying latency). This is the AI memory wall.
Compute Express Link (CXL) memory pooling offers a fundamentally different approach. Instead of tying memory to individual servers or GPUs, CXL creates a shared memory pool that multiple compute nodes can access dynamically, with latency that is orders of magnitude faster than network-based alternatives. For data center operators building and hosting AI training and inference infrastructure, CXL represents the most significant change to server architecture since the introduction of PCIe.
Understanding CXL: The Basics
Compute Express Link is an open industry standard interconnect built on the PCIe physical layer. It adds three protocols on top of the PCIe electrical interface:
- CXL.io: Standard PCIe I/O protocol for device discovery, configuration, and DMA operations.
- CXL.cache: Allows an attached device (like a GPU or accelerator) to cache data from the host CPU's memory with full hardware cache coherence.
- CXL.mem: Allows the host CPU to access memory attached to a CXL device using standard load/store instructions, as if it were local DRAM.
The CXL.mem protocol is what makes memory pooling possible. A pool of CXL-attached memory modules connected through a CXL switch becomes accessible to any compute node on that switch, with each node able to dynamically allocate and release memory capacity based on its current workload.
Key distinction: CXL memory access operates at 200-500 nanosecond latency, which is roughly 200-500x faster than NVMe storage access and only 2-4x slower than local DRAM. This makes CXL pooled memory viable for performance-sensitive AI workloads where NVMe offloading would be unacceptably slow.
CXL Specification Evolution
Understanding which CXL version enables which capability is essential for data center planning.
| CXL Version | Key Capability | Production Status |
|---|---|---|
| CXL 1.1 | Single-host memory expansion (Type 3 devices) | Shipping (2024+) |
| CXL 2.0 | Memory pooling with single-level switching | Shipping (2025+) |
| CXL 3.0/3.1 | Multi-level switching, fabric-attached memory, multi-host sharing | Early production (2026-2027) |
| CXL 4.0 | Multi-rack pooling, enhanced security, port flattening | Specification phase (2027+) |
For AI data center operators, CXL 3.0/3.1 is the inflection point. It enables the multi-host memory pooling and fabric-attached memory architectures that transform how GPU clusters are provisioned and operated.
How Memory Pooling Works
In a CXL memory-pooled architecture, the data center rack contains three categories of components:
- Compute nodes: Servers or GPU-equipped servers with CXL-capable CPUs and/or GPUs that connect to the CXL fabric.
- Memory modules: CXL Type 3 memory devices, essentially DRAM modules with CXL controllers, that sit in the shared pool. These can be standard DDR5 DIMMs behind a CXL controller or purpose-built CXL memory modules from vendors like Samsung, SK hynix, and Micron.
- CXL switches: Fabric switches that connect compute nodes to memory modules and manage address translation, access control, and quality of service. Marvell's Structera S 30260, with 260 CXL lanes, is among the first production-grade switches enabling rack-level pooling.
When a compute node needs more memory than its locally installed DRAM provides, it requests capacity from the pool through the CXL switch. The switch allocates a block of memory from an available module, maps it into the compute node's address space, and the node accesses it using standard memory instructions. When the workload completes, the memory is released back to the pool for allocation to other nodes.
Dynamic vs. Static Allocation
Early CXL pooling implementations use static allocation, where memory blocks are assigned to compute nodes at boot time or through management-plane configuration. This is simpler to implement and sufficient for many AI workloads where memory requirements are known in advance, such as hosting a specific large language model at a known batch size.
CXL 3.1 and beyond enable dynamic allocation, where memory can be assigned and released during runtime without node rebooting. This is critical for environments running variable workloads, such as an inference-as-a-service platform where different models with different memory footprints are loaded and unloaded throughout the day.
Why AI Workloads Need Disaggregated Memory
The economics of AI infrastructure make the case for memory pooling compelling.
The Stranded Memory Problem
In a conventional GPU cluster, each server is provisioned with enough DRAM to handle the peak memory requirement of any workload it might run. If a server hosts an NVIDIA GPU with 80 GB of HBM, the system might include 512 GB or 1 TB of CPU-attached DRAM to support data staging, preprocessing, and KV-cache overflow. But many workloads do not consume all that memory. Across a cluster, typical memory utilization in AI inference environments runs between 40 and 65 percent. The remaining 35 to 60 percent is stranded capacity: paid for but idle.
Memory pooling eliminates stranded capacity by making unused memory on one node available to other nodes. A 10-node cluster with 512 GB per node (5.12 TB total) running at 50 percent average utilization has 2.56 TB of stranded memory. With CXL pooling, that stranded capacity becomes available to nodes that need it, effectively increasing usable memory per node without buying additional hardware.
Scaling Model Size Without Scaling GPU Count
Without CXL, the only way to run a model that exceeds a single GPU's memory is model parallelism: splitting the model across multiple GPUs. This works but introduces inter-GPU communication overhead via InfiniBand or NVLink, adds scheduling complexity, and requires more GPUs per model instance.
CXL memory pooling offers an alternative path. By extending the addressable memory available to each GPU server through the CXL fabric, operators can keep models on fewer GPUs while offloading the excess memory footprint to pooled CXL memory at 200-500 ns latency. For inference workloads where the KV-cache is the primary memory consumer and is accessed less frequently than model weights, this tradeoff between slightly higher memory latency and significantly fewer GPUs can be economically favorable.
Right-Sizing Infrastructure
Disaggregated memory decouples compute procurement from memory procurement. Operators can buy GPU servers with minimal local DRAM and provision shared CXL memory independently, scaling each resource dimension to match actual workload requirements. This is analogous to how network-attached storage decoupled storage from compute, enabling more efficient resource utilization across the cluster.
CXL vs. Alternatives for Memory Extension
CXL is not the only technology for extending memory beyond local DRAM. Understanding how it compares to alternatives helps operators evaluate the right approach for their workloads.
| Technology | Latency | Bandwidth | Scope |
|---|---|---|---|
| Local DDR5 DRAM | ~80-100 ns | ~300 GB/s per socket | Single server |
| CXL-attached memory | ~200-500 ns | ~128 GB/s per link | Rack / multi-rack |
| InfiniBand RDMA | ~1-10 us | ~50 GB/s (400G) | Data center fabric |
| NVMe-oF (NVMe over Fabrics) | ~10-100 us | ~12 GB/s per device | Data center fabric |
| Host-managed SSD offload | ~100+ us | ~7 GB/s per device | Single server |
CXL occupies a unique position: fast enough for memory-semantic access (load/store instructions), with enough bandwidth for large-model AI workloads, but with a scope currently limited to rack-level fabrics. For AI inference workloads where KV-cache access patterns are bursty and latency-sensitive, the 200-500 ns CXL range is acceptable. For training workloads with sustained high-bandwidth memory access patterns, local HBM and DDR5 remain the performance tier, with CXL providing overflow capacity.
Data Center Infrastructure Implications
Deploying CXL memory pooling affects several aspects of data center infrastructure that operators and colocation providers must plan for.
Power and Cooling
CXL memory modules and switches consume additional power compared to a server-only deployment. A CXL switch handling rack-level pooling may draw 100-200 watts. CXL memory modules draw 5-15 watts each, comparable to standard DIMMs. The net power impact is modest relative to GPU server power consumption (which can reach 5-10 kW per server), but must be factored into rack-level power density planning and cooling capacity.
Rack Design and Cabling
CXL currently runs over PCIe 5.0 electrical connections, which limits cable length to approximately 2-3 meters for passive copper cables or up to 100 meters with active optical cables. This constrains memory pooling to within a rack or across adjacent racks. Future CXL specifications targeting multi-rack pooling will require optical interconnects, adding cost and cabling complexity to the infrastructure.
Management and Orchestration
CXL memory pools require management software to handle memory allocation, health monitoring, and quality-of-service policies. Integration with DCIM platforms and GPU orchestration tools like Kubernetes with GPU scheduling or Slurm is necessary for automated memory provisioning. The CXL Consortium's Fabric Manager specification defines a standard management interface, but vendor-specific tooling will vary in the early deployment phase.
Deployment Timeline and Planning
For data center operators planning AI infrastructure investments, CXL memory pooling is transitioning from technology preview to early production.
- 2025-2026: CXL 2.0 Type 3 memory expanders available for single-host memory expansion. Useful for extending per-server memory capacity without pooling. Early evaluation platforms from Samsung, Micron, and others.
- 2026-2027: CXL 3.1 switches and multi-host pooling reach early production. Marvell Structera and competitive products enable rack-level memory pools. Hyperscale early adopters begin production deployments.
- 2027-2028: Broader ecosystem maturity. OS, hypervisor, and orchestration software support stabilizes. Colocation providers begin offering CXL-ready rack configurations. Multi-rack fabric-attached memory becomes practical with optical CXL.
Operators deploying new high-density AI infrastructure today should consider CXL-readiness in their procurement decisions: selecting CPUs and motherboards with CXL support (Intel Sapphire Rapids and later, AMD Genoa and later), ensuring rack power budgets accommodate CXL switches and memory modules, and evaluating colocation partners on their roadmap for disaggregated infrastructure support.
How Rax Approaches Next-Generation AI Hosting Infrastructure
Rax designs its AI hosting infrastructure with the understanding that the boundary between compute and memory is dissolving. As CXL memory pooling matures from specification to production hardware, facilities that are designed with the power headroom, cabling infrastructure, and rack density to accommodate disaggregated architectures will serve tenants better than those built around fixed server-memory ratios. Rax tracks the CXL ecosystem closely and works with tenants to plan infrastructure that supports both current GPU cluster configurations and the disaggregated architectures that are coming next.
For operators evaluating managed AI hosting versus cloud GPU or colocation versus building their own, the flexibility to adopt CXL and other emerging interconnect technologies is a consideration that favors providers with the technical depth and infrastructure flexibility to support evolving hardware architectures.
Future-Ready AI Hosting Infrastructure
Rax operates GPU colocation and managed AI hosting in the UAE with infrastructure designed for current and next-generation hardware architectures. Talk to us about your AI compute and memory requirements.
Discuss Your Requirements