For the first few years of the AI boom, Nvidia was the undisputed leader in the market for state-of-the-art GPUs, reaping enormous profits as the artificial intelligence industry scaled up its compute capacity. The company's market cap grew roughly tenfold between the start of 2023 and mid-2025, fueled by an insatiable demand for its hardware. But in recent quarters, the narrative has shifted. Hyperscalers like Amazon, Google, and Microsoft have begun designing their own silicon, and investors have started to question how sustainable Nvidia's advantage truly is. The company's latest earnings report, however, suggests that the competitive picture is far more complex than a simple GPU race.
The new narrative centers around orchestration. As AI compute clusters grow to gigawatt scale, managing the flow of data and coordinating the numerous components of a data center has become a monumental challenge. Nvidia has quietly built a portfolio of specialized hardware and software designed to address precisely these challenges. This broader systems-level approach may prove more durable than GPU performance alone, and it is reshaping the competitive landscape in ways that many analysts are only beginning to appreciate.
Rack by rack
To understand this shift, it is worth examining exactly what Nvidia is selling today. The company is currently rolling out its Vera Rubin architecture, a comprehensive platform that pairs the next-generation Rubin GPU with a host of supporting components. These include the Vera CPU, a custom inference accelerator called Groq 3 LPX, as well as rack-level systems for storage and networking. None of these components are simple commodity parts; they are deeply specialized units designed to run in concert with the GPU and with each other.
The Vera CPU, in particular, is focused on the problem of orchestrating data movement. In a modern AI data center, the GPU is the workhorse, but it is dependent on a constant stream of data flowing from memory and storage systems. If that flow is disrupted or simply not managed efficiently, the GPU sits idle, wasting valuable compute cycles. "Vera is important because there's only so much memory that you can put in a single server or any sort of compute platform," said Jason Hardy, Nvidia's VP of storage technology, in a recent interview.
Hardy explained that as data centers have scaled up computing power, memory capacity has also grown, which has enriched companies like Micron during the second wave of the infrastructure boom. However, getting data to the GPU at the right time is not straightforward. As companies seek to drive tokens-per-watt (a key efficiency metric) lower, they are realizing that intelligent traffic direction is just as critical as raw processing power.
Nvidia claims that the Vera CPU produces dramatic improvements in these operations. "We saw upwards of 3x improvement in these operations, where the Vera CPU is allowing for acceleration," said Hardy. "So now we can use our flash to its fullest potential, because we can get all that performance out of it without bottlenecking."
The problem of data orchestration is not unique to Nvidia. When OpenAI developed its own custom chip, dubbed Jalapeño, a major design goal was avoiding these challenges altogether by minimizing the amount of data that needs to be moved. The company explained this in a blog post earlier this month: "We designed Jalapeño to minimize data movement and communication delays. Its large domain allows the entire workload to remain within one connected system, minimizing data movement and helping the complete request stay fast and efficient from beginning to end."
OpenAI's strategy is thus fundamentally different from Nvidia's. Rather than building specialized components to manage data flow across a sprawling architecture, OpenAI has chosen to integrate the workload into a single, large die. This reduces the need for orchestration in the first place. The underlying logic, however, is the same: efficiency increasingly depends on smarter traffic control rather than just additional processor cycles.
A new competitive arena
This convergence on data orchestration opens up a whole new layer of infrastructure for companies to compete over. For Nvidia, it is both an opportunity and a challenge. The company will have to defend its position against rival chipmakers and hyperscalers much as it has with GPUs. However, the competition has shifted to a different hierarchy: building a rival GPU now matters less than being able to make the entire system operate efficiently.
In the early stages, Nvidia appears to have a commanding lead. The company's long experience with GPU computing has given it deep insight into how data flows through a modern AI cluster. Moreover, its CUDA software ecosystem, while not the focus of this article, remains a powerful lock-in for developers. But the emergence of alternative approaches, such as Jalapeño, indicates that the future of AI hardware is not a one-horse race.
The hyperscalers are also not standing still. Amazon's Trainium and Inferentia chips, Google's TPUs, and Microsoft's Maia accelerators are all designed with similar goals of efficiency and integration. These efforts have already taken a toll on Nvidia's market perception, contributing to the more modest stock trajectory over the past year. Yet Nvidia's earnings report reveals a company that is still growing, still highly profitable, and arguably better positioned than ever to benefit from the next wave of AI infrastructure investment.
The gigawatt era
The shift to gigawatt-scale compute is one of the defining trends of the AI industry. While a typical enterprise server farm might consume a few megawatts, the largest AI training clusters now demand power in the range of hundreds of megawatts to a gigawatt or more. At that scale, even small inefficiencies in data movement, cooling, or power delivery can translate into enormous operational costs.
Nvidia has been adamant that its systems approach is ideally suited to this era. By designing the entire stack, from the GPU to the CPU to the networking and storage, Nvidia can tightly integrate every component to minimize waste. The Vera Rubin architecture is designed to be deployed at scale, with specialized racks and software that coordinate across thousands of nodes.
This is a significant departure from the early days of the AI boom, when companies could simply buy off-the-shelf GPUs and plug them into standard servers. Today, the complexity of AI workloads demands a more holistic approach. The winners in this new phase will be the companies that can deliver not just a chip, but a system that works as a coherent whole.
For Nvidia, this means competing on several fronts simultaneously. It must continue to advance the performance of its GPUs, while also ensuring that its surrounding architecture remains best-in-class. The company's investment in networking, through its acquisition of Mellanox, and its development of NVLink and InfiniBand technologies, underline its commitment to making the entire data center function as an integrated computing platform.
Data orchestration and the CPU's renewed relevance
One of the more unexpected consequences of the AI boom has been the renewed relevance of the CPU. For decades, the CPU was the heart of the computer, but the rise of GPUs relegated it to a supporting role. In the world of AI at scale, the CPU is making a comeback, not as a competitor to the GPU, but as the maestro that orchestrates the flow of data.
Nvidia's Vera CPU is a purpose-built product for this role. It is designed to handle the specific demands of AI workloads, including the need to rapidly move data between storage, memory, and GPU, while also supporting the complex scheduling and management tasks required in a mega-scale cluster. The fact that Nvidia has chosen to invest heavily in this component signals that it believes orchestration is where the next competitive battles will be won.
Hardy's comments about the 3x improvement in operations highlight the tangible benefits of this approach. By offloading orchestration tasks from the main compute path, the Vera CPU allows the GPU to focus on what it does best: running AI models. The result is higher overall efficiency, which in turn reduces the cost per token for operators and improves the economics of AI inference.
The competitive response
Nvidia's rivals are watching closely, and they are not without their own advantages. Hyperscalers, by virtue of controlling massive data center fleets, have a unique perspective on where the bottlenecks lie. They can design custom silicon that is tailored not only to their specific workloads but also to their own hardware and software stacks. This vertical integration is a powerful counter to Nvidia's horizontal platform approach.
OpenAI's Jalapeño chip is an example of a design that aims to transcend the need for orchestration altogether. By creating a large domain chip that can handle an entire workload without leaving the connectivity domain, OpenAI seeks to eliminate the delays that come from moving data off-chip. Whether this approach will scale to the largest training runs remains an open question, but it shows that the concept of data orchestration is being tackled from multiple directions.
For many industry observers, the real race is now between these competing philosophies. On one hand, Nvidia is building an increasingly complex ecosystem of specialized hardware and software. On the other hand, some of its biggest customers and potential rivals are looking to simplify the problem by integrating more functions into fewer, larger components. Both approaches have their merits, and the market will likely see a combination of the two for years to come.
What it means for the market
The stock market has already begun to adjust to this new reality. After the parabolic rise in shares through mid-2025, Nvidia's stock has been range-bound, as investors weigh the rise of custom silicon and the shifting competitive dynamics. The latest earnings report, however, helped to reassure the market that Nvidia's core business remains robust, and that the company is pivoting successfully to a new phase of growth.
One of the key takeaways from the earnings call was the emphasis on the full system, rather than just the GPU unit sales. Executives repeatedly highlighted the value of the complete rack-level solution, which commands a higher price point and better margins. This strategy also deepens customer lock-in, as the integration between components makes it harder for buyers to mix and match with third-party products.
There are risks, however. The increasing complexity of Nvidia's integrated platforms could also engender resistance from customers who prefer a more open ecosystem. The success of the hyperscaler custom chips suggests that many of the largest AI operators are eager to gain more control over their infrastructure. Nvidia's response is to make its systems so compelling that even the largest customers are reluctant to opt out.
Looking forward
As AI workloads continue to grow, the importance of orchestration will only increase. The era of simply adding more GPUs to a data center is giving way to a more sophisticated engineering challenge: how to keep those GPUs fed with data and coordinated with the rest of the system. Companies that master this challenge will be in a strong position to capture value from the next phase of the AI infrastructure buildout.
Nvidia's pivot to a systems approach is a calculated bet that this trend will play out in its favor. With its dominant GPU franchise, its comprehensive software stack, and now a suite of specialized hardware designed to optimize data movement, the company is aiming to remain the linchpin of AI infrastructure even as the competitive field becomes more crowded.
The road ahead is not without obstacles. The competition from hyperscalers and custom chip designers is real and intensifying. However, the company's early lead in the orchestration layer is a defense that was not fully appreciated just a few months ago. The market is beginning to recognize that Nvidia's advantage extends beyond the GPU - into the very fabric of the AI data center itself. Whether that advantage is sustainable will depend on how effectively the company can keep innovating in systems engineering as the era of gigawatt computing unfolds.
Source: TechCrunch News