
Analysts suggest that while Graphics Processing Units (GPUs) will maintain their dominance, the escalating corporate need for Artificial Intelligence (AI) inference is opening avenues for alternative solutions, according to Data Center Knowledge.
AI inference—the process of executing trained models to yield outputs—has become the industry’s new profit center. Major chip manufacturers are focusing on optimizing latency, power consumption, and cost. This is driving a shift toward a hybrid approach combining versatile GPUs with custom-designed silicon and attracting specialized engineering talent for their creation.
Nvidia’s $20 billion licensing agreement with Groq clearly signals this change. AMD acquired the engineering team at Untether AI, while Intel is reportedly contemplating the acquisition of SambaNova, currently valued around $1.6 billion. Analysts contend that this consolidation will not restrict the market; demand is expanding so rapidly that both established giants and emerging startups have room to compete within data centers and at the edge.
“Major corporations are strategically investing to bolster their inference solution portfolios, both through product development and talent acquisition,” states Matt Kimball, Vice President and Chief Analyst for Data Centers at Moor Insights & Strategy. “Numerous chip design startups are emerging, poised to make their mark.”
Why Inference is the Profit Center
According to Karl Freund, Founder and Principal Analyst at Cambrian AI Research, inference fundamentally differs from training in terms of both its economic structure and operational needs. While training AI models represents a cost center, inference is the “profit center” that directly generates revenue.
Freund and Kimball point out that although GPUs deliver outstanding performance, their architectural features are often geared towards training, which doesn’t always translate to lower latency or better efficiency for pure inference tasks. Specialized inference chips—ASICs and other accelerators—can offer faster response times, superior energy efficiency, and a reduced Total Cost of Ownership (TCO).
“From a profit center perspective, achieving low latency directly translates to increased revenue because users demand immediate responses, and you want to deliver them at the lowest possible expense,” Freund notes.
2026: Moving from Pilot Projects to Enterprise Production
Analysts observe that GPUs—led by Nvidia, with AMD gaining ground—currently dominate large-scale training and inference and will remain the choice for the largest workloads. However, the surging demand for inference is creating opportunities beyond GPUs, especially as major enterprises transition from pilot programs to full production deployment this year.
“We are seeing smaller organizations—perhaps those with 10,000 employees rather than 100,000—beginning to embed AI into their production, back-office, front-office, and edge processes,” says Kimball. These entities frequently encounter power constraints, cooling challenges, and ongoing GPU supply hurdles, rendering massive GPU cluster deployments impractical in many settings.
“When you deploy a GB200 or an H100, you are deploying something in the kilowatt range,” Kimball observes. “In retail settings, budgets for electricity are limited, and cooling infrastructure is poor, meaning you cannot utilize a GPU-heavy rack. They will seek alternative deployment solutions.”
For moderately sized organizations, such as a bank with 100 branches, the primary concerns simplify to TCO and energy budgets. This creates significant openings for inference-focused startups to meet their distinct requirements. “This is precisely where chip-design startups have substantial opportunity,” Kimball asserts. “It allows them to cater to customers who are currently underserved by existing market players, whether due to limited availability or specific power and performance demands.”
Inference Growth Spurs Market Diversification
Freund suggests that while GPUs remain the top general-purpose option for inference, the market is demonstrably leaning toward ASICs and alternative architectures provided by entities like AWS, Google, and various startups.
A Futurum Group survey conducted in November 2025 indicated that GPUs would account for 58% of data center compute spending by year-end, but in 2026, XPU architectures (neither GPU nor CPU) are projected to lead growth rates at 22%, surpassing GPUs (19%) and CPUs (14%).
“As the volume of inference tasks, measured in output tokens, surpasses the total volume of training tasks, there will be a strong imperative for diversification, given that alternative XPU architectures can deliver superior efficiency for certain defined inference workloads,” comments Brendan Burke, Futurum Group’s Director of Research for Semiconductors, Supply Chains, and Emerging Technologies.
Cloud Providers’ and Chip Suppliers’ Strategies
AWS’s approach reflects this growing requirement. The hyperscaler supports Nvidia, AMD, and Intel chips for AI workloads while concurrently offering its proprietary chips, giving customers choice, says Shaoen Nandi, AWS Director of Technology. He notes that while many clients prefer Nvidia for CUDA-optimized models, others are increasingly adopting AWS’s Trainium due to its performance-to-price ratio.
“Both options are highly sought after,” Nandi explains. “Nearly 50% of the tokens processed in Bedrock [AWS’s inference service] run on Trainium chips.”
Nvidia recognizes the necessity for specialized inference processors. Company executives reported that approximately 40% of its data center revenue in 2024 derived from inference. In September 2025, Nvidia unveiled the Rubin CPX, a GPU engineered for large-scale contextual inference in hyperscale and major enterprise environments, specifically targeting the pre-fill stage that handles the initial prompt before decoding. Nvidia’s licensing deal with Groq is reportedly aimed at embedding fast, low-latency, and cost-effective inference capabilities into its AI factory architecture; plans have been announced to leverage Groq’s low-latency processors to support broader real-time inference.
Intel is exploring multiple avenues to support inference beyond the proposed SambaNova acquisition. The company has enhanced its Xeon CPUs with AMX accelerators and offers dedicated Gaudi AI accelerators for inference workloads. “A significant portion of today’s inference resides on CPUs. A large portion of tomorrow’s inference will also run on CPUs,” Kimball states.
By securing the Untether AI engineering team, AMD further augmented its engineering capabilities with the November 2025 acquisition of MK1, an inference-specialized startup. MK1 is developing software that optimizes AMD GPUs for high-speed inference and reasoning in extensive enterprise deployments.
Freund predicts that Google’s latest TPU chip will pose a serious inference competitor, while Qualcomm’s upcoming AI200 and AI250 chips, which promise vast memory capacity and lower costs, could become appealing choices for data centers.
Startups to Monitor
Inference opportunities span data centers and the edge, with requirements varying significantly based on the workload and deployment location. “The inference you conduct inside your autonomous vehicle is vastly different from the inference you perform as an online customer support bot,” explains Kimball.
Jim McGregor, Principal Analyst at Tirias Research, highlights that inference capabilities are needed wherever computation occurs, including smartphones, PCs, and automobiles. “No two workloads are identical, and we will see a multitude of specialized AI accelerators designed for different load types,” he says. “This market is still nascent, with substantial room for many suppliers.”
Freund forecasts that in 2026, the majority of computation will still occur within data centers rather than at the edge.
Competitors in the data center inference space include Cerebras and Tenstorrent. Cerebras, founded in 2015, began offering inference capabilities on its chips about a year ago. According to Andy Hock, Senior Vice President of Products and Strategy at Cerebras, its systems can seamlessly switch between training and inference modes via software, with about 70% of their current workloads focused on inference; training still constitutes the bulk of the company’s revenue.
Tenstorrent, founded in 2016 and guided by industry veteran Jim Keller (who contributed to the AMD Zen architecture), is developing AI inference processors based on the RISC-V architecture.
South Korean NPUs are demonstrating diverse applicability, from edge devices to data centers. FuriosaAI is noted for its energy-efficient NPU architecture and has secured major clients like LG, Kimball points out. Rebellions, another South Korean startup, is recognized for its ARM-based technology and substantial backing from ARM and Samsung Ventures.