Scale up, Scale out, or Scale Across for AI Data Centers. Which is Best?16 min read

For years, there has been a discussion about whether it is best to scale up the capabilities of a server or rack or scale out to distribute the workload across multiple servers or racks.
Scaling UP within the rack or server means adding more CPUs, GPUs, RAM, and storage within one box. Known as vertical scaling, it centralizes and integrates everything in one machine, makes things easier to manage, and reduces communication delays. But there is only so much you can jam into one box or one rack. Additionally, scaling up has the liability of being a single point of failure.
Enter scaling OUT or horizontal scaling. This means adding more servers or racks in close proximity to share workloads. The idea behind this architecture is that it removes processing and storage limits from what can be included in one box. Provided it is supported by good networking, a more distributed approach offers better redundancy and simpler replacement of components and servers. On the downside, there is higher complexity to deal with, more possibility of networking bottlenecks, and integration headaches tend to multiply.
Scale up and scale out are both being implemented to accommodate the latest AI chips and servers. You can now pack far more into a rack than ever. Hence, rack densities are through the roof. But the pace of innovation has been so accelerated that scale up and scale out are now being supplemented by a third option: scale across.
Scaling ACROSS Data Centers
There is only so much processing that can be done in the latest GPU-filled racks. Yet the latest AI applications that seek to take “reasoning” to another level are now exceeding what can be done by multiple racks strung together and even what can be accommodated within one data center.
Enter scaling ACROSS, combining multiple data centers to act as one coordinated processing engine. According to Sameh Boujelbene, an analyst at Dell’Oro Group, as AI workloads are both compute-heavy and communication-heavy, the network is effectively being turned into one huge computer to tackle complex AI computations via scale across arrangements that set up multiple data centers to work in tandem.
Scale across, though, requires a significant upgrade of the network. The amount of network traffic will tend to compound as a great many highly dense racks within one data center are paired with dense racks in one or two other facilities. Boujelbene believes scale out AI will boost networking bandwidth requirements by roughly 10X, while scale across could push that closer to 100X.
Networking Vendors Respond
With scale out and scale across becoming a growing market for AI data centers, networking vendors have been rising to the challenge. The likes of Cisco, Marvell, Arista Networks, Nvidia, and Broadcom have been releasing new networking, switching, and routing products with greatly upgraded chips to prevent the network from becoming another AI bottleneck.
Specifically related to scale across, Cisco has released the P200 chip to boost networking connectivity between data centers running AI workloads. It is designed to shunt traffic efficiently between facilities at the highest speed possible on optical fiber. The P200 offers 51.2 Tbps of bandwidth as well as an external packet buffer that is packed with high-bandwidth memory. As a result, P200-based switches are designed to move AI-scale traffic across long distances on optical fiber while minimizing congestion and latency. This means that the fiber itself can operate close to maximum capacity for long periods, supported by a deep buffer to deal with congestion.
The Network as the Computer
AI is pushing the bounds of processing beyond what one rack, one aisle, or one data center can deliver. To match the performance of AI applications, it is increasingly about whether the network fabric can keep GPUs highly utilized, synchronized, and productive at a scale that spans multiple facilities.
Real-time monitoring, data-driven optimization.
Immersive software, innovative sensors and expert thermal services to monitor,
manage, and maximize the power and cooling infrastructure for critical
data center environments.
Real-time monitoring, data-driven optimization.
Immersive software, innovative sensors and expert thermal services to monitor, manage, and maximize the power and cooling infrastructure for critical data center environments.

Drew Robb
Writing and Editing Consultant and Contractor
0 Comments