...
.Nscale Locks $3.5 Billion Figure Robotics Compute Deal  ·Qatar’s Meeza Lands Major Hyperscaler Deal for 8MW ·Qualcomm Strikes Amazon AI Chip Deal, Opens Door to $4 Billion Stock ·Hitachi Energy Bets $300M on China Grid Manufacturing Corvex Builds Toward 8MW Cloud Infrastructure Footprint LITEON Bets $176 Million on DCX Liquid Cooling EdgeConneX Backs Singapore’s AI-Ready Tropical Data Center Testbed
.Nscale Locks $3.5 Billion Figure Robotics Compute Deal  ·Qatar’s Meeza Lands Major Hyperscaler Deal for 8MW ·Qualcomm Strikes Amazon AI Chip Deal, Opens Door to $4 Billion Stock ·Hitachi Energy Bets $300M on China Grid Manufacturing Corvex Builds Toward 8MW Cloud Infrastructure Footprint LITEON Bets $176 Million on DCX Liquid Cooling EdgeConneX Backs Singapore’s AI-Ready Tropical Data Center Testbed

The Neocloud Contract Problem Nobody Talks About: Cluster Fragmentation

A customer can have access to thousands of GPUs and still face a placement constraint. The problem emerges when a

Share
Neocloud Cluster Contracts

A customer can have access to thousands of GPUs and still face a placement constraint. The problem emerges when a distributed workload needs resources inside a specific topology domain. Aggregate free capacity does not guarantee enough suitable resources within that boundary. Topology-aware schedulers account for this distinction when they place large jobs. Slurm can organize nodes into contiguous blocks and allocate jobs according to topology. Kubernetes also supports topology-aware placement for grouped workloads. Large AI jobs often require communication across many accelerators during execution. That communication makes network placement part of the available compute environment. A contract based only on GPU quantity may leave this requirement undefined.

Reserved GPUs Are Not Necessarily a Reserved Cluster

Capacity agreements can encourage buyers to think primarily in GPUs, nodes, instances, or accelerator-hours. Distributed AI software sees a more complex resource. Workers exchange gradients, activations, parameters, and other data while a job runs. Network paths between those workers can matter alongside accelerator specifications. This becomes particularly relevant for communication-intensive distributed training. NVIDIA reference architectures combine compute, networking, storage, and cluster-management components within an integrated design. They also distinguish different network functions across the infrastructure. Two nodes with identical GPUs can therefore occupy different positions in the network hierarchy. An agreement based only on aggregate quantity may not define those placement differences.

Fragmentation Starts With Placement, Not Failure

Cluster fragmentation does not require a hardware failure. It can emerge when free resources become scattered across topology domains. The total GPU count may still look sufficient for the requested workload. Yet those GPUs may not form an acceptable allocation under the scheduler’s topology rules. Slurm addresses this issue through topology-aware allocation models. Its block approach can group nodes and preserve useful placement structures. Smaller jobs can occupy available portions while larger jobs wait for a feasible allocation. Scheduler configuration, priority, topology requirements, and cluster state all influence that outcome. The infrastructure can remain operational even when a specific distributed job cannot obtain its required placement.

How a Healthy Cluster Becomes Commercially Fragmented

Fragmentation can develop through normal scheduling activity rather than one major infrastructure event. Jobs start and finish at different times. Nodes may leave the available pool during maintenance. Reservations can also overlap with other scheduled activity. These events create openings in different parts of the topology. Schedulers must decide how to use those resources while serving current requests. Slurm supports allocation strategies that consider switches, blocks, and other topology relationships. Those mechanisms show why placement efficiency involves more than counting idle nodes. Aggregate capacity can remain available while the resources suitable for one workload change.

A buyer may have limited visibility into the placement policies behind a hosted GPU service. That matters when a workload needs a particular network or topology relationship. Shared infrastructure can add another scheduling dimension because active jobs occupy different resource combinations. Their allocations can affect the shapes available to later requests. This does not mean shared infrastructure will always create fragmentation. The result depends on architecture, scheduler policy, workload size, and current cluster state. Customers should identify whether placement characteristics matter to their applications before signing a reservation. Providers can then define which characteristics form part of the service. That creates a clearer distinction between raw inventory and workload-ready capacity.

Why Network Domains Change the Meaning of Availability

Physical network structure determines how distributed workers communicate. Placement can therefore change the communication environment available to a job. Current NVIDIA rack-scale designs provide a clear example. NVLink supplies scale-up connectivity within a rack-scale domain. InfiniBand can provide scale-out connectivity between racks in those documented systems. Expanding beyond one domain can therefore introduce a different communication layer. Application effects depend on workload design, parallelism, software, and network configuration. No universal performance penalty should be assumed from topology alone. Therefore, buyers should distinguish accelerator quantity from the placement characteristics supporting those accelerators.

The distinction becomes important when customers benchmark one configuration and later receive another. Identical GPU models do not guarantee identical network placement. A workload might also cross more topology boundaries as its allocation grows. That change does not automatically make the new allocation unsuitable. It does create a technical variable that buyers may need to test. Representative benchmarks can reveal whether placement materially changes application behavior. Contract terms can then focus on properties that actually affect the workload. This avoids demanding unnecessary locality while protecting requirements that have measurable value. Capacity becomes easier to evaluate when its physical shape is no longer implicit.

The Contract Can Guarantee Quantity While Leaving Shape Undefined

Compute capacity can be described through accelerator type, memory, node configuration, quantity, and service duration. Topology-sensitive workloads introduce another consideration: where those resources sit when scheduled. A buyer might need homogeneous nodes within an acceptable network domain. The same number of nodes scattered elsewhere may not satisfy that requirement. Scheduler documentation confirms that topology can influence allocation decisions. Jobs can receive placement according to switch, block, or other infrastructure relationships. An agreement based on accelerator quantity alone may not address those relationships. Whether that omission has contractual significance depends on the specific agreement. Procurement teams should identify which infrastructure properties matter before treating reserved quantity as deployable capacity.

Contiguity Needs a Technical Definition

Simply requesting “contiguous capacity” does not create a precise technical requirement. Contiguity can mean different things across cluster architectures. It might describe GPUs inside one server or nodes inside a scale-up domain. Another design could define locality around switches, racks, or allocation blocks. NVIDIA architectures contain several distinct connectivity layers. Slurm also supports configurable topology structures rather than one universal cluster layout. Buyers should translate workload needs into measurable infrastructure attributes. These could include accelerator type, node homogeneity, interconnect class, or maximum placement scope. The specification should follow the workload rather than impose unnecessary architecture.

Scheduling Policy Becomes Part of the Product

Customers purchase compute capacity, but scheduler behavior influences when that capacity becomes usable. Slurm illustrates how allocation policies can account for locality and fragmentation. Kubernetes also supports topology-aware scheduling for grouped workloads. Its current topology-aware functionality remains an evolving capability rather than a universal hosted-cluster feature. The broader engineering principle is more important than the scheduler brand. Some distributed workloads require resource relationships that a simple node count cannot describe. Slurm and Kubernetes can both represent topology during placement. Individual providers may use different orchestration systems and policies. Customers need to understand the resulting service behavior rather than every internal scheduler detail.

Small Jobs Can Affect Large Reservations

Fragmentation creates a tension between immediate utilization and future placement options. Small jobs can fit into openings that larger allocations cannot use individually. Slurm’s block topology illustrates this resource-management problem. Its design allows smaller jobs to use available space within partially occupied blocks. The scheduler can also prioritize placement that reduces fragmentation across the cluster. Cluster operators therefore have several placement objectives to balance. However, the appropriate balance depends on the service and scheduling model. Dedicated capacity can behave differently from shared or on-demand pools. Buyers should not infer large-job placement guarantees from headline accelerator counts alone.

Homogeneity Adds Another Fragmentation Layer

Physical proximity does not make every accelerator operationally equivalent. Nodes can differ across hardware and software characteristics. Possible differences include GPU generation, memory configuration, drivers, firmware, and network interfaces. Kubernetes exposes GPUs through device plugins and supports hardware-aware placement through labels and selectors. Those mechanisms allow workloads to request nodes with required characteristics. A distributed job may therefore impose several constraints at once. It may need enough GPUs, suitable topology, compatible hardware, and a matching software environment. Every additional requirement narrows the resources eligible for that allocation. Free accelerator inventory should not automatically be interpreted as deployable capacity for every workload.

GPU Sharing Changes What a Resource Count Can Represent

GPU sharing adds another reason to define capacity precisely. NVIDIA’s GPU Operator supports time-slicing of physical accelerators. Administrators can expose multiple schedulable replicas backed by the same underlying GPU. Those replicas share the physical device. They do not provide the same memory or fault isolation associated with MIG. NVIDIA also warns that requesting several time-sliced replicas does not guarantee proportional compute power. A scheduler-visible resource count can therefore represent different underlying arrangements. Moreover, buyers need to know whether capacity means physical GPUs, partitions, or shared scheduler resources. That definition should come before any discussion about topology or fragmentation.

Expansion Can Expose Fragmentation That Smaller Jobs Hide

A cluster can serve smaller jobs successfully and encounter new constraints when those jobs expand. Larger requests need more resources to become available at the same time. They may also need those resources within specific topology boundaries. Block-oriented scheduling exists partly to preserve useful resource groupings for such allocations. Customers planning workload growth should examine more than future GPU quantity. They should also test whether expanded allocations retain important infrastructure characteristics. Crossing racks or network domains can alter communication paths without changing the GPU model. The performance impact requires benchmarking and cannot be predicted from topology alone. Scaling contracted devices does not automatically produce an identical increase in immediately schedulable distributed compute.

Maintenance Can Alter the Available Shape

Maintenance can temporarily reduce the resources available inside a topology domain. Whether that creates a placement constraint depends on the remaining cluster state. Scheduler configuration, workload size, and topology requirements also matter. Capacity elsewhere in the cluster may remain usable for other workloads. Yet a topology-constrained job might still lack a feasible placement. Kubernetes documents this general scheduling condition for grouped topology-aware workloads. NVIDIA architectures also illustrate why replacement location can affect the available network relationship. Buyers can ask how maintenance affects placement characteristics promised by their service. Contract terms can define which technical properties should remain consistent when underlying resources change.

Buyers Need a Schedulable-Capacity View

Procurement teams can reduce ambiguity by connecting inventory definitions with workload placement. Accelerator model, quantity, memory, and node configuration remain important. Buyers should also understand how those resources can be grouped. Relevant questions include topology boundaries and constraints on large allocations. Scheduling behavior deserves attention when the service uses shared resources. Customers can ask what happens when a requested allocation cannot be placed immediately. These questions do not require disclosure of another tenant’s information. They require enough technical detail to assess whether the purchased service matches the workload. Instead, relying only on aggregate counts can leave important topology requirements unresolved.

Technical diligence can close much of that gap before deployment. Buyers can run representative workloads against the proposed infrastructure. Architecture disclosures can clarify which network domains an allocation may cross. Service terms can document placement requirements that testing identifies as important. The objective is not to turn every infrastructure detail into a contractual guarantee. It is to identify properties that materially affect usable compute. A raw GPU count remains valuable, but it describes only one resource dimension. Placement rules determine which combinations can satisfy certain distributed jobs. Procurement becomes more precise when both measurements appear in the evaluation.

Placement Guarantees Should Match Workload Requirements

Not every AI workload needs strict topology guarantees. Smaller independent inference replicas may tolerate broader placement than tightly coupled training jobs. Even training workloads can have very different communication profiles. Model architecture, parallelization strategy, batch configuration, framework behavior, and scale can all matter. Buyers should avoid demanding maximum locality without evidence that the application needs it. Representative benchmarking can identify which placement properties materially affect performance. Those findings can become acceptance criteria where contractual protection is justified. Providers can distinguish guaranteed properties from best-effort characteristics. This approach protects workload requirements without imposing unnecessary scheduling restrictions.

Observability Has to Extend Beyond GPU Utilization

Allocated GPU count and utilization provide only part of the operational picture. Topology-sensitive customers may also need information about scheduling outcomes. Useful telemetry can include pending-job reasons, queue duration, and eligible resource counts. Some platforms may expose topology or placement information as well. Available fields depend on the scheduler and provider. Slurm uses topology information during allocation. Kubernetes can also coordinate grouped workloads around topology constraints. Where suitable telemetry exists, customers can distinguish placement problems from other scheduling delays. That visibility makes service reviews more useful than a simple comparison of installed and allocated GPUs.

Observability also helps separate several problems that can look similar from the application layer. A job might wait because resources are busy. Another job could remain pending because eligible resources exist in the wrong topology domains. Configuration requirements can further reduce the candidate node set. These conditions can require different operational responses. A utilization dashboard alone may not reveal those distinctions. Customers do not need unrestricted access to the provider’s internal scheduler. They need enough information to understand the state of their contracted workloads. Clear status information also creates a better foundation for SLA measurement.

The SLA Needs a Response to Fragmentation

A topology-aware service needs a defined response when an agreed resource shape cannot be assembled. The mechanism should reflect the workload and provider’s service model. Options can include placement windows, queue thresholds, alternate topology, or capacity substitution procedures. Commercial remedies may also apply where the parties negotiate them. These mechanisms are contract choices rather than universal industry standards. Ordinary infrastructure uptime does not necessarily show whether a topology-constrained allocation can run. Hardware can remain operational while a specific workload lacks a feasible placement. Fragmentation-aware schedulers demonstrate that this is a genuine resource-management condition. Clear service thresholds can address it without promising instantaneous placement for every possible workload.

Capacity Planning Must Account for the Shape of Free Space

High accelerator utilization and immediate availability for large contiguous jobs are different scheduling objectives. A cluster can have substantial free capacity distributed across several locations. That free space may still fail to satisfy one topology-constrained request. Preserving larger blocks can influence how smaller workloads receive placement. Slurm’s block topology explicitly addresses fragmentation when allocating resources. The mechanism demonstrates why the shape of free capacity matters alongside its quantity. A topology-specific commitment also differs from an aggregate accelerator commitment. The first restricts which resources can satisfy the customer’s allocation. Buyers should evaluate those service characteristics alongside price and nominal GPU quantity.

This distinction also changes how buyers should compare apparently similar offers. Two providers can advertise the same accelerator generation and device count. Their scheduling models, network domains, and placement commitments can still differ. Those differences do not automatically make one service better. The appropriate configuration depends on the customer’s workload requirements. A loosely coupled application may value scheduling flexibility more than strict locality. Large distributed training can place greater importance on predictable resource relationships. Technical evaluation should identify which characteristics affect the intended workload. Commercial comparison can then focus on capacity that the customer can actually use.

The Contract Should Define Compute as a System

The lesson is not that neocloud capacity guarantees are inherently unreliable. Nor does every distributed workload suffer from fragmentation. Accelerator quantity simply represents one dimension of an AI compute service. Current architectures connect GPUs through structured scale-up and scale-out networks. Schedulers can use topology information when deciding where workloads run. Those layers affect which combinations of hardware can satisfy particular requests. A contract can reflect this reality without prescribing every infrastructure detail. It can define the resource characteristics, placement expectations, substitution rules, and visibility that matter. That gives both parties a clearer definition of what reserved compute means.

For buyers, this approach changes the question asked during procurement. “How many GPUs are reserved?” remains necessary, but it is no longer sufficient for every workload. Teams should also ask what resource shape can be scheduled under agreed conditions. They should identify the topology boundaries that matter to tested application behavior. Providers can then state which requirements they can support consistently. That process turns an ambiguous capacity promise into a more measurable service definition. It also separates genuine workload requirements from preferences that add little technical value. Cluster fragmentation becomes manageable when both sides treat schedulable topology as part of capacity. The contract then describes not just the hardware purchased, but the compute configuration the workload can actually consume.

[simple-author-box]

More from AI Infrastructure

Artificial-intelligence infrastructure is running into an increasingly unusual constraint: Electricity can exist without being

The US House has passed a bipartisan bill that could reshape how America assigns

A neocloud announcing another campus or GPU cluster can sound like straightforward good news

COMPUTE WEEKLY

The briefing that 40,000+ tech leaders read every Monday. Sharp, fast, essential.

Building an AI Startup Without Owning GPUs

Not owning GPUs has become the default, deliberate strategy for building an AI company — not a compromise founders accept reluctantly. H100 rental rates fell 64-75% in fifteen months, a dense ecosystem of neoclouds and inference-as-a-service providers now lets startups skip infrastructure entirely, and credit programs can fund a company’s first year before a founder writes a check
Most Read

A compute node sitting behind a garage door can perform the same basic computational

A project can leave a site without leaving behind the conditions that made the

A commercial operation date can look precise long before the underlying project is capable

A 5 GW AI infrastructure plan can satisfy every conventional site-selection requirement and still

A fire strategy becomes expensive when the building has already decided where walls, equipment,

Disruptor Spotlight

Cerebras Systems

The chip that makes Nvidia nervous. Cerebras’ Wafer Scale Engine is rewriting the rules of AI inference at scale.
Faster
0 x
YoY Revenue
0 x
Transistors
0 T
Market Pulse
MSFT
+1.02%
NVDA
+0.66%
AMZN
-0.078%
AMD
-6.95%
TSMC
-2.98%
Indicative only · Not financial advice
Upcoming Events
SEP
The AI Infrastructure Race (India)
WEBINAR · ONLINE
The AI Infrastructure Race: Won on Power, Land and Trust — Not Capital
MAY
0
AI Infrastructure Summit
DUBAI · IN PERSON
MEA’s premier AI infrastructure event.
JUN
0 0
Compute Forecast Summit
SINGAPORE · IN PERSON
Our flagship APAC event. Early bird open.
Latest Moves
Live
ecolab
Ecolab Deepens Cooling Strategy With $4.75B CoolIT Acquisition
Ecolab is making one of its biggest moves yet into AI infrastructure after completing its $4.75 billion acquisition of liquid cooling specialist CoolIT Systems
Pure DC AVK Europe data center microgrid Dublin 110MW AI infrastructure Ireland 2026
Pure DC and AVK Deploy Europe’s First 110 MW Data Center Microgrid in Dublin
The Pure DC Dublin microgrid has made history as Europe’s first large-scale on-site data center microgrid, launched in partnership with power solutions provider AVK at Pure DC’s campus in Ireland.
Pace Digitek
Pace Digitek Partners With MEGMEET to Expand AI Data Center Power Business
India’s AI infrastructure ecosystem continues to mature as domestic technology manufacturers move beyond traditional telecommunications and industrial markets toward high-growth digital infrastructure opportunities
Follow Compute Forecast
11K followers
1200 followers
Companies to Watch
CW
CoreWeave
Neo Cloud · $19B · IPO Watch
CB
Cerebras Systems
AI Hardware · $4.25B · Pre-IPO
G42
G42
Sovereign AI · Abu Dhabi
H
Humain
Saudi AI · $40B Fund
Latest Podcast
AI Capex, Cloud Margins & the Nuclear Bet
48 MIN · 25 APR 2026

The Neocloud Contract Problem Nobody Talks About: Cluster Fragmentation

A customer can have access to thousands of GPUs and still face a placement constraint. The problem emerges when a

Share
Neocloud Cluster Contracts
0
847 SHARES

0
SHARES

[simple-author-box]

More from AI Infrastructure

A compute node sitting behind a garage door can perform the same basic computational

A project can leave a site without leaving behind the conditions that made the

A commercial operation date can look precise long before the underlying project is capable

A 5 GW AI infrastructure plan can satisfy every conventional site-selection requirement and still

COMPUTE WEEKLY

The briefing that 40,000+ tech leaders read every Monday. Sharp, fast, essential.

Great! We’ve received your information.

Global AI Infrastructure Outlook 2026

The briefing that 40,000+ tech leaders read every Monday. Sharp, fast, essential.
Download Free
Most Read

A compute node sitting behind a garage door can perform the same basic computational

A project can leave a site without leaving behind the conditions that made the

A commercial operation date can look precise long before the underlying project is capable

A 5 GW AI infrastructure plan can satisfy every conventional site-selection requirement and still

A fire strategy becomes expensive when the building has already decided where walls, equipment,

Disruptor Spotlight

Cerebras Systems

The chip that makes Nvidia nervous. Cerebras’ Wafer Scale Engine is rewriting the rules of AI inference at scale.
Faster
0 x
YoY Revenue
0 x
Transistors
0 T
Market Pulse
NVDA
$924.60
+2.4%
MSFT
$421.30
+1.1%
AMZN
$192.80
-0.6%
NVDA
$924.60
+2.4%
NVDA
$924.60
+2.4%
Indicative only · Not financial advice
Upcoming Events
MAY
0 0
DCD Global — London
LONDON · IN PERSON
World’s largest DC event. CF is media partner.
MAY
0
AI Infrastructure Summit
DUBAI · IN PERSON
MEA’s premier AI infrastructure event.
JUN
0 0

Compute Forecast Summit

SINGAPORE · IN PERSON
Our flagship APAC event. Early bird open.
Latest Moves
  • Live
Sam Altman
OpenAI appoints new Chief Infrastructure Officer to lead $100B DC programme
27 APR · OPENAI
Sam Altman
OpenAI appoints new Chief Infrastructure Officer to lead $100B DC programme
27 APR · OPENAI
Sam Altman
OpenAI appoints new Chief Infrastructure Officer to lead $100B DC programme
27 APR · OPENAI
Follow Compute Forecast
18.4K followers
12.1K followers
9.3K subscribers
41 episodes
Companies to Watch
CW
CoreWeave
Neo Cloud · $19B · IPO Watch
CB
Cerebras Systems
AI Hardware · $4.25B · Pre-IPO
G42
G42
Sovereign AI · Abu Dhabi
CW
Humain
Saudi AI · $40B Fund
Latest Podcast
AI Capex, Cloud Margins & the Nuclear Bet
48 MIN · 25 APR 2026
Scroll to Top
Seraphinite AcceleratorOptimized by Seraphinite Accelerator
Turns on site high speed to be attractive for people and search engines.