Building AI infrastructure is no longer just about buying chips; it is a sequential engineering challenge that starts with the grid. A recent deep-dive into the physical stack reveals that power availability often dictates the pace and ultimate size of an AI campus. Before a single GPU is installed, developers must navigate generation, transmission, and substations, each representing a distinct stage of project maturity. The distinction between contracted power and energized capacity is critical for anyone planning large-scale deployments.
The Power Bottleneck
NVIDIA has proposed a significant shift toward bringing 800-volt direct current (DC) closer to the rack. This architecture aims to reduce the number of power conversion stages, thereby minimizing energy waste and creating physical space for additional computing equipment. The transition is described as a move from current AC facilities through hybrid systems to native 800 VDC sites. However, it is crucial to note that most operating data centers do not use this architecture today, making it a forward-looking optimization rather than an immediate standard.
Cooling and Facility Constraints
The facility layer is equally constrained by thermal physics. Almost every watt consumed by computing equipment converts to heat, meaning the cooling system strictly limits rack density and runtime. The largest current facility records highlight the scale of this challenge; for instance, the Colossus 2 campus is estimated at 946 MW, a draw comparable to a mid-sized city. Other major projects, including Amazonβs New Carlisle site at 910 MW and Microsoftβs Fairwater Atlanta at 636 MW, illustrate the massive energy footprint required for modern AI clusters.
From Tokens to Tasks
Inside the rack, the division of labor between GPUs and CPUs is becoming more complex. While GPUs handle parallel math measured in tokens, CPUs manage sequential tasks like tool calls and browser sessions. NVIDIAβs Vera CPU, featuring 88 custom Olympus cores, is purpose-built for this agentic workload. The article emphasizes that public benchmarks for complete agent workflows are only beginning to emerge, leaving CPU performance metrics like 'tasks per second' less standardized than GPU token throughput. Developers must carefully size host CPUs to prevent expensive GPUs from idling while waiting for sequential tool calls to complete.
Key Takeaways
- NVIDIAβs 800 VDC proposal reduces conversion losses but is not yet standard in most operating data centers.
- Cooling systems, not just power supply, are the primary limiting factor for rack density in large AI campuses.
- The Vera CPU and Blackwell GPU pairing highlights a shift toward measuring performance in 'tasks' for agentic workflows, distinct from traditional token metrics.
- Facility records show massive energy demands, with top-tier campuses approaching 1 GW of estimated power draw.
The Bottom Line
The physical stack is the new bottleneck for AI scaling. While NVIDIA's 800V DC and Vera CPU innovations promise efficiency gains, the reality is that most infrastructure remains stuck in legacy AC and thermal constraints, forcing builders to prioritize grid access and cooling capacity over raw chip acquisition.