Scale-beyond — the next shift in AI infrastructure

Home Analyst Angle Scale-beyond — the next shift in AI infrastructure
AI grid AI infrastructure scale-beyond

As the center of gravity shifts from training to inference, the telecoms industry has an opportunity to provide distributed AI infrastructure and capture new revenue

The AI infrastructure buildout has so far centered on massive, centralized data centers optimized for training. But as the market shifts toward inference and toward latency-sensitive use cases such as physical AI, the defining challenge is becoming less about concentrating compute and more about coordinating it across a distributed, heterogeneous environment.

That was the central theme of a recent webinar featuring Nokia Leader for AI and High Performance Networking Clayton Wagar, Red Hat’s Atul Deshpande, chief principal architect in the Field CTO Office for the global telco vertical, and CoreWeave Vice President of Product Management for AI Platforms and Services Urvashi Chowdhary. Together, the speakers described an emerging AI fabric in which connectivity, orchestration and cloud platforms work across organizational and geographic boundaries to place workloads according to performance, sovereignty, availability and cost.

Watch the full webinar here.

Wagar framed Nokia’s “scale-beyond” concept as the next step in AI infrastructure after scale-up, scale-out and scale-across architectures. Whereas those models expand a single compute domain, scale-beyond connects independent AI domains into a larger system. The resulting network will need higher capacity, deterministic performance, real-time telemetry, closed-loop automation and security embedded at every layer.

“The connectivity fabric of the AI grid is scale beyond,” Wagar said.

The implication is that networks cannot remain passive transport. AI workloads will be increasingly symmetrical, machine-generated and continuous, while placement decisions may change dynamically based on latency, energy availability, regulation or capacity. Workload-management systems this distributed AI infrastructure plant will therefore need to signal requirements into the network so bandwidth, resiliency, routing and policy can be adjusted programmatically.

Deshpande focused on the orchestration layer required to make that distributed model operational. Generative AI workloads differ from conventional cloud-native applications because prompts and responses are non-uniform, accelerators are costly, and inference is stateful and benefits from cache-aware routing. Red Hat’s approach combines an AI gateway, distributed inference, optimized models and a common operational layer spanning multiple accelerator families and hybrid-cloud environments.

“The portability of the stack … across different hardware architectures — that’s probably one of the key decisions,” Deshpande said. He also highlighted production evidence that state-aware routing can increase capacity and reduce time to first token without adding hardware, reinforcing the importance of software-level optimization.

Chowdhary brought the discussion back to customer outcomes. Training is a controlled, tightly coupled job; inference introduces bursty demand, request-level service requirements, multi-tenancy and the need to scale down between spikes. That makes workload placement and unit economics central design considerations.

Her architectural principle was direct: “Open at the interfaces and integrated underneath.” Platforms should absorb provisioning, failure recovery and operational toil, she said, while preserving enough control for customers to tune hardware selection, topology, quantization and price-performance. The ultimate metric should not be “useful work per dollar.”

The shared conclusion was that AI networking is shifting from a capacity problem to a coordination problem. Success will depend on understanding each workload, preserving portability and governance, and measuring whether placement decisions improve real business outcomes rather than simply moving compute closer to the edge.

What you need to know in 5 minutes

Join 37,000+ professionals receiving the AI Infrastructure Daily Newsletter

This field is for validation purposes and should be left unchanged.

This website uses cookies to improve your experience. We'll assume you're ok with this, but you can opt-out if you wish. Accept Read More