AI networking is reaching a point where knowing where a packet should go is no longer enough to decide how the fabric should behave. A switch can forward traffic correctly while still making the wrong decision for the computation that depends on that traffic, particularly when several distributed jobs compete for the same paths, queues, and links. The important question therefore shifts from whether the packet can reach its destination to whether the network understands the workload relationship behind the packet and can preserve that relationship while conditions change. Cisco has increasingly described this direction through context-aware traffic arbitration, identity-aware forwarding, and network intelligence embedded closer to the switching path, while NVIDIA has built Spectrum-X around congestion control, adaptive routing, and performance isolation for AI traffic.
Traditional switching begins with a comparatively simple proposition: inspect the packet, determine the relevant forwarding information, select a path, and move the packet toward its destination. That model remains fundamental, but AI clusters expose a limitation because the forwarding decision says little about why the traffic exists or how its timing relates to other traffic generated by the same computation. A distributed training job can create communication patterns in which many GPU endpoints exchange data as part of a coordinated operation, making congestion on one path materially different from congestion generated by an unrelated workload. Identity-aware forwarding gives the fabric an additional context layer by incorporating traffic identity and visibility into the forwarding architecture rather than relying only on conventional destination information.
Why Identity Becomes a Forwarding Primitive
The architectural significance of identity appears when the switch must distinguish between traffic that looks similar at the packet level but has very different consequences at the workload level. Two flows can traverse the same physical fabric and consume comparable resources while belonging to jobs with completely different communication patterns, service expectations, and tolerance for delay. A destination-based forwarding table cannot, by itself, express that difference because its central purpose remains determining reachability rather than interpreting computational intent. Identity-aware forwarding introduces the possibility of carrying contextual information into the decision process so that traffic treatment can reflect the workload relationship behind the flow. Cisco’s description of identity-aware forwarding in Silicon One points toward this evolution by placing identity and traffic visibility directly into the silicon rather than leaving all contextual interpretation to a separate operational layer.
That distinction can also change how operators think about policy because identity-aware forwarding can provide visibility into the ‘who, what, and where’ of network traffic at each node, giving operators contextual information beyond destination reachability. A workload can move across hosts, containers can be recreated, and application endpoints can receive different addresses without changing the underlying role of the computation. In an AI fabric, the same principle can support more consistent treatment when a distributed job spans many compute nodes and its traffic pattern changes during execution. The network does not need to infer intent from an address alone if the workload context can remain associated with the traffic through the relevant enforcement and forwarding mechanisms.
From Visibility to Actionable Fabric Context
Visibility becomes useful only when the network can translate what it observes into an operational decision. AI fabrics therefore need more than counters showing queue occupancy, link utilization, or packet movement because those measurements describe network conditions without necessarily explaining which computation created the pressure. NVIDIA’s Spectrum-X architecture uses telemetry, congestion control, adaptive routing, and performance isolation as coordinated mechanisms, illustrating how network observations can influence traffic treatment instead of remaining passive operational data. Its approach also recognizes that AI traffic requires the fabric to respond to congestion at a finer level than traditional static routing can provide.
Cisco’s current AI networking direction presents a related movement toward workload-aware telemetry and policy control, with its Nexus networking material describing visibility into AI traffic and workload behavior as part of the fabric architecture. The important architectural point is that telemetry becomes more valuable when the network can associate observed behavior with a workload rather than merely reporting that a port or queue is busy. That association can support more precise traffic steering, policy enforcement, and troubleshooting because the operator can reason about the computational source of a network condition. It also creates a path toward automation in which the fabric reacts to workload behavior without requiring every decision to pass through an external management system.
Context-Aware Congestion Management at Scale
Congestion in an AI fabric is not merely a queue that has become full because too much traffic arrived at once. The queue represents the visible result of competing communication patterns, and those patterns can belong to workloads with different dependencies, priorities, and sensitivity to disruption. A conventional congestion response can reduce pressure by slowing traffic, redirecting flows, or applying queue-based controls, but those actions do not inherently distinguish whether the affected traffic belongs to one critical distributed computation or several unrelated background operations. NVIDIA’s Spectrum-X design explicitly combines telemetry-based congestion control with performance isolation so that multiple AI workloads can share the network without one workload’s traffic creating disproportionate disruption for another.
A workload-aware fabric can also make isolation more precise because it can separate the identity of the traffic source from the physical location where congestion appears. This distinction becomes important when a workload’s traffic crosses multiple switches, since a local queue condition can originate from a broader communication pattern that spans the fabric. Cisco’s workload-aware networking direction emphasizes real-time network insights and dynamic congestion isolation, while its identity-aware forwarding work places richer traffic context inside the networking silicon. Together, those capabilities point toward an architecture where congestion can be evaluated through traffic conditions, telemetry, routing state, and forwarding context rather than treated solely as an anonymous condition attached to a single interface.
Selective Backpressure Instead of Fabric-Wide Throttling
Backpressure becomes more sophisticated when the fabric can identify the traffic responsible for a condition and understand which traffic should remain protected. Without that context, an aggressive response can spread congestion control too broadly, reducing useful throughput for workloads that did not create the original contention. AI fabrics need a more selective model because different jobs can coexist on the same physical network while requiring different treatment during transient pressure. NVIDIA describes Spectrum-X performance isolation as a coordinated combination of quality-of-service isolation, adaptive routing, and congestion control, which provides a technical example of how the network can limit interference without simply shutting down competing traffic.
Selective response also requires the switch to make decisions quickly enough to remain relevant to the traffic condition. An external controller can understand a broader system state, but a control loop that depends on repeated telemetry collection, interpretation, and configuration changes can operate at a different timescale from packet forwarding. Moving more intelligence into the switching path reduces that distance between observation and action, allowing the fabric to respond while the congestion event is still active. Cisco’s description of identity-aware forwarding embedded in Silicon One illustrates this movement toward silicon-level context, while NVIDIA’s adaptive routing and congestion-control mechanisms show a parallel effort to make real-time traffic decisions within the network itself.
Performance Isolation Without Stranding Capacity
Performance isolation becomes difficult when several AI workloads share the same physical fabric because strict separation and efficient utilization can pull the architecture in opposite directions. A completely dedicated network path can simplify contention management, but it can also leave capacity unused whenever the associated workload does not generate traffic. A completely shared network can improve utilization, yet uncontrolled contention can allow one communication pattern to influence the progress of another. The engineering challenge therefore sits between physical dedication and unrestricted sharing, requiring the fabric to isolate harmful interference while continuing to exploit resources that remain available. NVIDIA positions Spectrum-X around this problem by combining congestion management, adaptive routing, and mechanisms designed to maintain predictable performance across concurrent AI workloads.
Isolation works differently when the network recognizes the workload rather than treating every flow as an equivalent consumer of bandwidth. A workload can generate many flows across multiple paths, and those flows can interact with traffic from other jobs at several points inside the fabric. If the network applies policy independently to each flow without understanding that relationship, it can protect individual queues while failing to protect the computational behavior that those queues collectively support. Traffic-aware policy allows the fabric to apply differentiated treatment to traffic classes and use mechanisms such as congestion control, adaptive routing, and performance isolation to manage interference among concurrent AI workloads. Cisco’s networking work around identity-aware forwarding reflects this movement by bringing identity into the forwarding decision rather than leaving workload context entirely outside the switching path.
Protecting Performance Without Reserving Everything
The most useful isolation mechanism therefore does not attempt to create a private network for every workload. It creates a policy boundary that the shared fabric can enforce dynamically while preserving the ability to use spare resources elsewhere. That distinction lets the network treat isolation as a behavioral guarantee rather than a physical reservation, which becomes increasingly important as workloads vary in communication intensity. Spectrum-X describes performance isolation as a core capability for multi-tenant AI environments, with the network using congestion control and adaptive routing to limit the effect of one workload on another. A switch that knows workload identity can also make isolation decisions with greater precision than a system that sees only interfaces and queues. Traffic belonging to one job can receive differentiated treatment without forcing unrelated traffic into the same restriction simply because it happens to traverse the same physical resource.
The consequence is a shift from capacity reservation toward capacity arbitration. Reservation assumes that protection requires setting aside resources before traffic arrives, whereas arbitration assumes that the fabric can observe contention and decide how resources should be distributed as conditions change. AI networking benefits from the latter approach because distributed workloads can change their communication behavior during execution, making fixed allocations less responsive to actual demand. The switch consequently becomes responsible for protecting workload relationships without turning every shared resource into a dedicated resource. That is the foundation for a fabric in which performance isolation can coexist with high utilization rather than forcing operators to choose one at the expense of the other.
In-Fabric Scheduling: When Arbitration Moves Into Silicon
The advantage of in-fabric arbitration comes from reducing the distance between observation and response. A controller can collect information from across the network and build a broader picture, but the switch sees the immediate state of its own queues, links, and forwarding choices at the moment traffic arrives. That local perspective allows the switching system to make rapid forwarding and path-selection decisions using current network conditions, while adaptive routing and congestion-control mechanisms operate directly within the networking stack. NVIDIA’s adaptive routing and congestion-control mechanisms demonstrate how traffic decisions can move into the network fabric itself rather than depending solely on centralized coordination.
Identity makes this local arbitration more meaningful because the switch can attach context to the traffic it is already processing. A packet does not need to become a complete representation of application intent for the network to benefit from knowing which workload or traffic class it belongs to. The switching system can combine forwarding context with network conditions to support rapid path selection, traffic management, and congestion response. Cisco’s work is important in this respect because Cisco explicitly describes identity-aware forwarding, real-time telemetry, and deep traffic visibility as capabilities embedded directly in Silicon One.nThis architecture also changes the relationship between the control plane and the data plane. The control plane can establish identities, policies, topology information, and broader intent, while the data plane can execute rapid decisions using the context that has already been established.
Scheduling as a Fabric-Native Function
Once arbitration becomes part of the switching pipeline, traffic scheduling and resource arbitration increasingly operate as integrated functions of the forwarding architecture, with switching systems using queueing, traffic classification, routing, and congestion-management mechanisms to control how traffic accesses network resources. The switch does not merely determine the next hop; it determines how traffic competes for the available forwarding resources under current conditions. That can include choosing among viable paths, managing queues, responding to congestion signals, and maintaining isolation between traffic classes whose behavior should not interfere with each other. NVIDIA’s description of Spectrum-X places these functions within a unified Ethernet architecture designed specifically around the communication characteristics of AI workloads.
Silicon-level scheduling also creates a requirement for carefully bounded intelligence because every additional decision consumes hardware resources and can affect deterministic packet processing. The objective cannot simply be to insert more policy into the ASIC; the objective must be to identify which decisions provide meaningful workload-level benefit and can execute within the timing and resource constraints of the forwarding pipeline. Identity-aware forwarding therefore becomes valuable when it supplies concise context that can influence a decision without requiring the switch to reconstruct the entire application state. Cisco’s work around identity and telemetry points toward precisely this kind of selective context, where the network becomes more aware without becoming an application execution environment.
The Transition to Intent-Driven Fabric Architecture
Aggregate throughput remains necessary because an AI fabric cannot compensate for inadequate physical capacity through intelligence alone. The more consequential question, however, concerns what the network does with that capacity when several workloads compete for it at the same time. A fabric can deliver strong link utilization and still create inefficient computation if congestion management ignores the relationship between traffic and workload progress. Identity-aware forwarding adds traffic identity and visibility to the forwarding architecture, while AI-oriented congestion-control, adaptive-routing, and performance-isolation mechanisms provide additional ways to manage shared network resources under contention. Cisco’s workload-aware networking direction and NVIDIA’s Spectrum-X architecture both support the broader conclusion that AI networking increasingly depends on coordinated treatment of traffic rather than forwarding performance considered in isolation.
Intent-driven networking does not mean that the switch receives an abstract instruction such as “make this workload faster” and independently discovers how to accomplish it. The more realistic model uses explicit identities, policies, traffic classifications, congestion signals, and routing information that collectively express what the fabric should protect and how it should respond. Those inputs allow the network to combine forwarding information with traffic identity, telemetry, routing state, and congestion conditions while keeping immediate decisions within the established functions of packet forwarding and network resource management. Cisco’s identity-aware forwarding work provides an example of how identity can become part of that decision model, while NVIDIA’s Spectrum-X architecture demonstrates how congestion control, routing, and isolation can work together around AI traffic behavior.


