...
.Nscale Locks $3.5 Billion Figure Robotics Compute Deal  ·Qatar’s Meeza Lands Major Hyperscaler Deal for 8MW ·Qualcomm Strikes Amazon AI Chip Deal, Opens Door to $4 Billion Stock ·Hitachi Energy Bets $300M on China Grid Manufacturing Corvex Builds Toward 8MW Cloud Infrastructure Footprint LITEON Bets $176 Million on DCX Liquid Cooling EdgeConneX Backs Singapore’s AI-Ready Tropical Data Center Testbed
.Nscale Locks $3.5 Billion Figure Robotics Compute Deal  ·Qatar’s Meeza Lands Major Hyperscaler Deal for 8MW ·Qualcomm Strikes Amazon AI Chip Deal, Opens Door to $4 Billion Stock ·Hitachi Energy Bets $300M on China Grid Manufacturing Corvex Builds Toward 8MW Cloud Infrastructure Footprint LITEON Bets $176 Million on DCX Liquid Cooling EdgeConneX Backs Singapore’s AI-Ready Tropical Data Center Testbed

Your Next GPU Upgrade May Actually Be a Cooling Upgrade

Why the Next GPU Upgrade Starts With Cooling GPU performance often leads the conversation when an organisation plans to expand

Share
GPU Cooling

Why the Next GPU Upgrade Starts With Cooling

GPU performance often leads the conversation when an organisation plans to expand its artificial intelligence infrastructure. Training workloads may require more computing power, while inference services must handle growing demand. A newer accelerator can deliver higher performance, more memory or better performance per watt. However, these advantages do not automatically increase useful output from an existing data centre. The surrounding infrastructure must supply enough electrical power, remove the resulting heat and maintain suitable operating conditions. A rack designed for older accelerators may lack the cooling distribution, electrical delivery or physical space required by a denser replacement system. Even when the facility has sufficient total power, individual racks or cooling zones can reach their practical limits first. This creates a planning challenge because hardware procurement can advance while facility readiness remains uncertain. The key question is not simply whether a newer GPU fits inside the existing rack. It is whether the complete infrastructure can support that hardware under its expected operating conditions. Cooling therefore deserves attention alongside accelerator selection, rather than as a facilities upgrade that follows the purchase.

Service Public Policy Newsletter Leaderboard 970x118 1

How Higher GPU Density Changes the Cooling Equation

A server upgrade changes more than the number of calculations a system can perform. Accelerator power, memory configuration, networking and packaging also influence the heat that a facility must remove. High-performance GPU servers can concentrate substantial electrical demand within a relatively small rack footprint. This creates local thermal loads that differ from those of conventional enterprise servers. The challenge involves both the total heat load and the rate at which cooling equipment can transfer heat away from processors and other components. Air cooling remains suitable for many data centre workloads, but its practical limits depend on airflow, equipment design, inlet temperatures and installation density. High-density systems may require direct liquid cooling, close-coupled cooling or a combination of liquid and air systems. The right choice depends on the server design and its thermal requirements. Direct liquid cooling transfers heat from selected components through circulating liquid, often using cold plates attached to processors or other high-power devices. However, other components may still rely on airflow. Liquid cooling does not automatically eliminate the need for conventional cooling equipment.

Rack Power Is Only the Starting Point

A rack’s electrical rating provides a useful starting point for assessing infrastructure requirements. However, it does not capture every condition that determines reliable operation. Two racks with similar power consumption may distribute heat differently because their server layouts, cooling interfaces and airflow paths differ. A system that transfers most processor heat into liquid still generates heat from memory, power supplies, storage, networking and other components. The facility must remove these remaining loads through a suitable combination of air and liquid infrastructure. Engineers also need to examine cooling performance during sustained operation. A short commissioning test or theoretical equipment rating cannot establish how the system will perform under every operating condition. Coolant temperature, flow rate, pressure drop and heat-exchanger performance all influence heat removal. The equipment manufacturer’s operating requirements provide the necessary reference for evaluating the proposed cooling arrangement. Ignoring these details can lead teams to overestimate the usable compute capacity of a facility.

When Air Cooling Stops Being the Simplest Option

Air cooling remains a proven approach for conventional server rooms and lower-density racks. It supports a wide range of equipment without requiring liquid connections at every server. However, its suitability depends on the combined heat load, airflow management, room design and environmental limits of the installed hardware. As rack power increases, moving enough air through the equipment becomes more difficult. This challenge becomes particularly important when server fans, rack layouts and containment systems do not provide clear paths for hot and cold air. Higher fan speeds can increase energy consumption and acoustic output without resolving every local hotspot. Close-coupled systems and rear-door heat exchangers can help some facilities manage higher loads without converting every rack to direct liquid cooling. Their effectiveness depends on the amount of heat they can remove, the equipment’s operating limits and the facility’s wider heat-rejection capability. A design that works for one high-density rack may not suit an entire cluster with different layouts or load profiles. Therefore, cooling architecture should follow measured heat loads and equipment requirements rather than a universal density threshold.

Service Advisory Services Leaderboard 970x118 1

Direct Liquid Cooling Introduces a Wider Infrastructure Requirement

Direct liquid cooling moves heat away from high-power components through a circulating liquid in a properly designed circuit. A cold-plate system typically moves coolant through plates attached to selected components. The system then transfers the absorbed heat through a cooling distribution unit or another suitable interface to the facility infrastructure. This arrangement introduces pumps, pipes, valves, sensors, connectors and controls that require correct installation and ongoing maintenance. The facility must also match coolant supply temperatures and flow conditions to the IT equipment’s requirements. Liquid cooling therefore changes the engineering boundary between the server and the building. It involves more than replacing a fan with a pipe. The provider and hardware operator need clear responsibilities for coolant supply, connection maintenance and responses to alarms or leaks. Commissioning should verify the complete cooling path, including the interface between the technology cooling loop and the facility system.

Cooling Readiness Must Be Verified Before Hardware Arrives

A procurement team can confirm a GPU’s technical specifications and delivery date without establishing whether the intended data hall supports the final server configuration. This gap matters because the hardware, electrical distribution, cooling equipment and network fabric must all be ready before the cluster can deliver its planned workload. A facility assessment should begin with the manufacturer’s rack-level power and thermal requirements. Engineers can then review available electrical capacity, cooling distribution and heat-rejection capability. The assessment should identify whether the installation requires direct liquid cooling, supplemental air cooling or modifications to existing equipment. Engineers should also check the intended rack position for suitable pipe routes, service clearances, floor loading capacity and installation access. A site-wide megawatt figure cannot confirm that every row or room has the same usable capacity. This is especially important when a facility contains several cooling zones with different capabilities. The review should document constraints that could affect deployment and identify who must resolve each issue. Procurement decisions should therefore rely on verified site readiness. The hardware delivery date alone cannot determine when deployment can begin.

The Cooling System Can Limit GPU Utilisation

Installing more powerful accelerators does not guarantee sustained performance throughout a long training run or demanding inference session. GPU performance depends on several variables, including workload characteristics, power limits, processor temperatures, memory behaviour and software configuration. Thermal conditions matter because processors operate within manufacturer-defined limits. Their control systems can adjust operating behaviour when temperatures or power conditions require intervention. The precise response varies by device and workload, so teams should not attribute every performance shortfall to cooling. Infrastructure teams need telemetry that connects GPU temperatures and utilisation with rack power, coolant conditions and facility operating data. These measurements help engineers distinguish thermal constraints from software bottlenecks, memory pressure, network congestion or insufficient workload parallelism. A sustained performance decline may justify examining the cooling path. However, diagnosis should rely on correlated measurements rather than a single temperature reading. The goal is to establish whether the infrastructure can maintain the required operating conditions at the workload’s intended scale.

Measure the Complete Thermal Path

A useful monitoring plan should connect device-level measurements with the cooling equipment responsible for removing rack heat. GPU temperature and power readings can reveal changes in accelerator behaviour. Rack inlet and exhaust temperatures provide additional context about the surrounding airflow. Liquid-cooled systems also require suitable measurements of supply and return temperatures, flow rates and relevant pressure conditions. A coolant temperature within its specified range does not prove that every connected component receives sufficient flow. Engineers must interpret these readings alongside workload demand, pump operation, valve positions and equipment-specific limits. The monitoring platform should distinguish measurements from the server, the rack and the facility cooling system. This separation helps operators understand where a developing problem originates. Correlating these data sources can reveal restrictions before they become deployment or availability problems.

Cooling Efficiency Matters Alongside Cooling Capacity

A cooling system can remove the required heat without operating at the lowest possible energy consumption. Pumps, fans, chillers, heat exchangers and other mechanical equipment all contribute to facility energy use. Their combined demand varies with climate, system design, operating temperatures and workload. An upgrade should therefore consider both the amount of heat the facility can remove and the energy required to do so. Some liquid-cooling designs support higher coolant temperatures. Where site conditions and equipment specifications permit, these temperatures can reduce reliance on mechanical refrigeration. Other installations require lower supply temperatures or additional refrigeration, changing the efficiency calculation. Cooling equipment selection must account for local ambient conditions, water availability where relevant, maintenance requirements and the intended operating envelope. A system that performs well in one climate or facility layout may deliver different results elsewhere. Energy modelling and operational measurements can help teams compare alternatives. These assessments should not assume that every liquid-cooled installation automatically consumes less energy.

Avoid Treating Liquid Cooling as a Universal Efficiency Guarantee

Liquid cooling offers effective heat transfer, but its overall energy and operational benefits depend on the complete system. A design may reduce some server fan requirements while adding pump demand, coolant distribution equipment and additional controls. Facility-side energy use also depends on the heat-rejection method and operating temperatures. A retrofit that introduces liquid cooling into an existing data hall may require additional heat exchangers or changes to the chilled-water arrangement. The engineering team should model these requirements against the expected IT load. It should then compare the results with the current system’s performance. Operational measurements after commissioning can establish whether the upgrade delivers the anticipated cooling capacity and energy results. The review should also consider maintenance access, spare components and the consequences of losing a pump, valve or cooling distribution unit. A credible business case must distinguish verified improvements from design assumptions that still require testing.

Retrofitting an Existing Data Centre Requires a Wider Review

An existing facility may have enough unused electrical capacity for a new GPU cluster but still require mechanical or physical upgrades. The review should examine rack power distribution, cooling capacity, pipe routes, floor loading, equipment clearances and maintenance access before teams approve the final layout. Some liquid-cooling installations can use existing facility cooling infrastructure through a suitable cooling distribution unit. However, compatibility depends on the selected system’s temperature, flow and pressure requirements. A facility may also need to retain air cooling for components that transfer their heat into the liquid circuit. Structural considerations deserve attention because high-density racks, coolant distribution equipment and associated pipework can introduce loads that differ from older server installations. The proposed arrangement should preserve access for component replacement and equipment isolation during maintenance. Meanwhile, the project team should coordinate hardware delivery with construction, commissioning and operational readiness. This approach reduces the risk of receiving expensive accelerators before the installation area can support them.

Commissioning Should Test the Intended GPU Configuration

Commissioning must validate the system that operators intend to run, rather than simply confirm that individual components switch on. The test plan should use the relevant hardware configuration, cooling interface and operating conditions. This helps establish whether the installation meets its specified requirements. Engineers should verify coolant connections, leak detection where applicable, flow conditions, temperature limits, alarms and control behaviour. Air-cooled components also need appropriate checks because liquid-cooled racks may still rely on airflow to remove residual heat. The commissioning process should test responses to defined failure scenarios, including equipment loss where the design requires redundancy. Test results should document the operating conditions and any limitations that remain after initial validation. A controlled workload can help teams correlate thermal measurements with actual GPU activity. However, testing must follow the hardware manufacturer’s procedures. The final acceptance record should identify unresolved issues, operating limits and the party responsible for closing each outstanding item.

Cooling Planning Should Shape the Next GPU Procurement Cycle

GPU procurement decisions often involve performance targets, memory requirements, network compatibility, software support and delivery schedules. Cooling should form part of the same assessment because it can determine which rack configurations a facility can deploy. It also affects the additional infrastructure an installation may require. Procurement teams should request the intended system’s power and cooling specifications early enough for facility engineers to assess compatibility. This review should happen before purchase commitments become difficult to change. Teams must distinguish capacity that is already available from capacity that depends on new electrical or mechanical work. They should also establish whether the existing facility supports the required liquid-cooling interface or needs additional equipment and commissioning. A cost comparison should include infrastructure modifications, installation labour, energy use and maintenance responsibilities. It should also account for any deployment delays caused by facility upgrades. Finally, the team should confirm that the selected configuration fits the facility’s physical and operational limits. A nominal power allocation alone cannot establish that the complete system is ready. This approach makes the procurement decision more representative of the total cost and readiness of the compute system.

The Upgrade Decision Must Account for the Entire System

The value of a GPU upgrade depends on the performance that the surrounding infrastructure can support, not just the accelerator specification. A deployment may require changes to cooling distribution, electrical delivery, network capacity, rack layout and monitoring before it operates as intended. The balance varies between a small inference deployment, a dense training cluster and a mixed-use data hall. Teams should assess the actual workload instead of applying a single infrastructure template. A facility with suitable air cooling may not need a full liquid-cooling conversion. A high-density configuration, however, may require direct liquid cooling alongside complementary air systems. The engineering assessment should identify the constraints that matter for the selected hardware and estimate the work required to address them. Operational telemetry can then validate whether the installed system maintains the required conditions under sustained workloads. Procurement, facilities and compute teams should use consistent assumptions when estimating deployment schedules and total cost of ownership. The next accelerator investment will be stronger when teams assess the cooling design, facility readiness and intended workload as one connected engineering decision.

Service Podcast Leaderboard 970x118 1
[simple-author-box]

More from AI Infrastructure

Advanced Nuclear Power Takes On AI’s Energy Hunger

The rapid expansion of artificial intelligence is forcing utilities to reconsider an old assumption:

The Transformer Replacement Timeline Is Becoming an AI Resilience Question

Why Transformer Planning Is Moving Into AI Strategy AI data center transformer planning is

Singapore Enforces Stricter Security Standards for Major Data Centers

Singapore has turned data center resilience into a matter of national infrastructure policy, moving

COMPUTE WEEKLY

The briefing that 40,000+ tech leaders read every Monday. Sharp, fast, essential.

Building an AI Startup Without Owning GPUs

Not owning GPUs has become the default, deliberate strategy for building an AI company — not a compromise founders accept reluctantly. H100 rental rates fell 64-75% in fifteen months, a dense ecosystem of neoclouds and inference-as-a-service providers now lets startups skip infrastructure entirely, and credit programs can fund a company’s first year before a founder writes a check
Most Read

The Data Hall Is Only One Part of Deployment A data center can have

A data center design can look complete on paper while remaining impossible to build

A tower can put a surprisingly small amount of computing space inside a surprisingly

Cooling failure rarely arrives at the compliance desk as a clean regulatory event, because

Ireland’s experience with data-center expansion became less about stopping construction than about changing the

Disruptor Spotlight

Cerebras Systems

The chip that makes Nvidia nervous. Cerebras’ Wafer Scale Engine is rewriting the rules of AI inference at scale.
Faster
0 x
YoY Revenue
0 x
Transistors
0 T
Market Pulse
MSFT
+1.02%
NVDA
+0.66%
AMZN
-0.078%
AMD
-6.95%
TSMC
-2.98%
Indicative only · Not financial advice
Upcoming Events

TBC

The AI Infrastructure Race

WEBINAR · ONLINE
The AI Infrastructure Race: Won on Power, Land and Trust — Not Capital
MAY
0
AI Infrastructure Summit
DUBAI · IN PERSON
MEA’s premier AI infrastructure event.
JUN
0 0
Compute Forecast Summit
SINGAPORE · IN PERSON
Our flagship APAC event. Early bird open.
Latest Moves
Live
Ecolab
Ecolab Deepens Cooling Strategy With $4.75B CoolIT Acquisition
Ecolab is making one of its biggest moves yet into AI infrastructure after completing its $4.75 billion acquisition of liquid cooling specialist CoolIT Systems
Gemini Generated Image 5gy41q5gy41q5gy4
Pure DC and AVK Deploy Europe’s First 110 MW Data Center Microgrid in Dublin
The Pure DC Dublin microgrid has made history as Europe’s first large-scale on-site data center microgrid, launched in partnership with power solutions provider AVK at Pure DC’s campus in Ireland.
Clipboard Image 1784558387
Pace Digitek Partners With MEGMEET to Expand AI Data Center Power Business
India’s AI infrastructure ecosystem continues to mature as domestic technology manufacturers move beyond traditional telecommunications and industrial markets toward high-growth digital infrastructure opportunities
Follow Compute Forecast
11K followers
1200 followers
Companies to Watch
CW
CoreWeave
Neo Cloud · $19B · IPO Watch
CB
Cerebras Systems
AI Hardware · $4.25B · Pre-IPO
G42
G42
Sovereign AI · Abu Dhabi
H
Humain
Saudi AI · $40B Fund
Latest Podcast
AI Capex, Cloud Margins & the Nuclear Bet
48 MIN · 25 APR 2026
Scroll to Top
Seraphinite AcceleratorOptimized by Seraphinite Accelerator
Turns on site high speed to be attractive for people and search engines.