Enterprise technology leaders are increasingly abandoning the idea of purchasing individual hardware components in favor of integrated AI systems. At the recent AMD Advancing AI event, analysts emphasized that the real value of artificial intelligence now comes from how these systems are architected and how they support actual business workflows, rather than the isolated accelerators used in the past. This shift marks a move away from fragmented technology stacks toward full-stack AI infrastructure that unifies compute, networking, memory, and software.
Dave Vellante, co-CEO of SiliconANGLE Media Inc. and co-host of theCUBE Research, noted that the current situation suggests AMD does not need to displace Nvidia to succeed. The company aims to become an essential second player in the AI market. This strategy relies on AMD’s massive investments in mergers and acquisitions, specifically the $49 billion acquisition of Xilinx, which seeded the ecosystem with software like ROCm. The goal is to provide enterprises with a viable alternative that offers flexibility and control.
Bob O’Donnell, president at TECHnalysis, pointed out that AMD’s software strategy is key for this transition. By evolving ROCm to allow GPU programs to be rewritten from CUDA into ROCm format, the company is helping to reduce the barriers for developers looking to move off Nvidia’s dominant platform. This evolution of software is becoming a key factor in how long-term platform openness is evaluated by enterprise IT departments.
Related: Nscale Acquires Anyscale for 1 Billion Dollars
Flexible deployment across environments
Enterprises are no longer asking for specific CPUs or GPUs, but rather solutions that address concrete business challenges. Derek Dicker, corporate vice president of the Enterprise Business Group at AMD, explained that customers want integrated strategies that combine the best-of-breed technology available. This requires infrastructure that supports consistent software across cloud, hybrid, and on-premises environments, allowing organizations to maintain flexibility without facing vendor lock-in.
Operational strategy is also evolving to balance different types of models. Suresh Andani, CVP of compute and enterprise AI at AMD, stated that companies are increasingly running frontier models in the cloud while hosting open-weight models on-premises. This approach allows organizations to optimize for cost, performance, and governance simultaneously. It is a central topic in nearly every conversation regarding AI deployment, as providers seek to enable efficient hosting of these open-weight models on local infrastructure.
While the focus is often on the silicon, the physical implementation of these systems requires deep integration. Microsoft’s collaboration with AMD spans everything from data center facilities and power distribution to networking and software. Alistair Speirs, GM of Azure infrastructure at Microsoft, highlighted that these systems must be designed as a whole unit, rather than as separate pieces. This holistic approach ensures that rack-scale AI infrastructure can handle the demands of high availability and reliability.
Related: Man jailed due to username mistake
As AI workloads scale, the networking layer becomes a foundational element of rack-scale infrastructure. AMD and Meta Platforms Inc. have co-designed the networking architecture behind Helios to ensure that the Ultra Accelerator Link over Ethernet can effectively connect all GPUs. Soni Jiandani, senior VP and GM at AMD, noted that this design addresses the complexity of failure domains, allowing dozens of GPUs to act as a single, high-memory unit.
Operational visibility across these distributed environments is also critical. Cisco Systems Inc. contributes a management layer through its Cloud Control platform, which allows administrators to view inferencing capacity in the cloud, private data centers, and endpoint devices from a single management plane. Jeetu Patel, president and chief product officer at Cisco, explained that this unified view is necessary to manage the growing complexity of modern AI deployments effectively.
