AI Semiconductor/GPU Economics: NVIDIA / TPU / Trainium

AI Navigate Original / 4/27/2026

💬 OpinionSignals & Early TrendsIdeas & Deep AnalysisIndustry & Market Moves
共有:

Key Points

  • AI competition is decided by compute procurement at huge scale
  • NVIDIA dominant; challengers: TPU, Trainium, Maia, Groq/Cerebras ASICs
  • China uses Ascend/SMIC under export controls
  • Bottlenecks: TSMC/HBM/power/cooling; multi-cloud/chip strategy

AI Competition's Essence Is Chip Procurement

The AI race is decided not just by how smart a model is, but by "how much compute you can secure, and how cheaply." OpenAI, Anthropic, Google, Meta, and xAI lock up GPUs and in-house accelerators at enormous scale, and governments fund sovereign-AI compute programs of their own.

NVIDIA's Dominance and Challengers

NVIDIA: Hopper → Blackwell → Rubin

NVIDIA's line has moved Hopper (H100/H200) → Blackwell (B200 / GB200 NVL72) → Rubin, and the "Vera Rubin NVL72" rack unveiled at CES 2026 is now in full production on TSMC 3nm. NVIDIA claims up to 5x the inference performance of Blackwell and up to a 10x lower cost per token, with broad adoption expected from the second half of 2026. A high-end GPU still costs tens of thousands of dollars and a full rack can reach roughly 100M yen, and demand keeps outrunning supply, so multi-month lead times are normal. NVIDIA's strength is not the GPU alone but the whole box: the NVLink interconnect plus the CUDA software base.

Google TPU

The current TPU is the 7th generation, "Ironwood" (TPU v7), which carries 192GB of HBM per chip and links thousands of chips into a single pod. It backs Gemini training and inference and is also offered externally through GCP (Anthropic and various startups).

AWS Trainium / Inferentia

Amazon's AI-dedicated chips. The current training generation, Trainium3, is built on TSMC 3nm with 144GB of HBM3e per chip, and "Project Rainier" — the giant cluster whose main customer is Anthropic — now totals on the order of a million Trainium-family chips. The aim is training and inference that costs less than NVIDIA.

Microsoft Maia / Cobalt

Microsoft's in-house AI chips for Azure. Used to run OpenAI models and to serve Azure customers, with the aim of reducing NVIDIA dependence and operating cost.

Dedicated Inference ASICs

  • Groq LPU: dedicated to ultra-low-latency inference; it sells on how quickly the first tokens come back
  • Cerebras WSE: a giant design that turns one whole wafer into a single chip, so a large model need not be split across devices
  • SambaNova: enterprise-dedicated
  • Etched Sohu: a Transformer-dedicated ASIC that trades generality for efficiency
  • AMD Instinct MI400 / Helios: AMD's rack-scale answer, due in the second half of 2026 at 72 GPUs and 31TB of HBM4 per rack — the first serious rack-for-rack challenger to NVIDIA's Vera Rubin

China's Situation

US export controls kept NVIDIA's top GPUs out of China and NVIDIA sold regulation-compliant versions instead. In July 2026 the US side eased those controls and Beijing was reported to be approving H200 imports for domestic AI companies (SCMP, GIGAZINE), reopening a channel that had been limited to cut-down parts. China's own response:

  • Huawei Ascend (910C): the flagship domestic AI chip, with roughly 600,000 units planned for 2026
  • The real ceiling is HBM: domestic high-bandwidth memory supply is expected to cover only about 300,000 chips' worth in 2026, which caps how many accelerators China can actually finish
  • Domestic manufacturing rests on SMIC (7nm) and domestic HBM (CXMT); startups such as Biren and Moore Threads are also funded
  • DeepSeek and others show that strong models can still be built under the constraints

Bottlenecks

BottleneckStatus
GPU manufacturingTSMC 5nm/3nm monopoly. CoWoS packaging tight
HBM memory3-firm oligopoly: SK Hynix, Samsung, Micron
PowerData-center power orders-of-magnitude; nuclear/renewable/self-gen
Liquid coolingMandatory in Blackwell gen, huge facility-retrofit cost
LocationCompetition for land suited to transmission/water/cooling

What to Watch in 2026

  • How fast the Blackwell → Rubin transition runs: Rubin volume production started before Blackwell supply had settled, so long-term GPU contracts are being reopened
  • How far Google TPU v7 Ironwood and AWS Trainium3 can actually stand in for NVIDIA
  • Whether AMD MI400 / Helios becomes a real rack-for-rack alternative
  • Whether inference-dedicated chips such as Cerebras win production workloads on throughput
  • HBM and CoWoS tightness plus data-center power, which set the ceiling for everyone
  • How far Huawei can push past its HBM constraint
  • OpenAI achieved its US AI-compute-infrastructure 10GW target several years ahead of schedule (April 2026)
  • NVIDIA shipped "Vera," its first CPU designed for AI agents, to Anthropic, OpenAI, SpaceXAI, and Oracle Cloud (a dedicated CPU now joins the GPU-centric landscape)

Implications from a Corporate View

  • Secure cloud GPU prices with long-term contracts
  • Compare inference cost across AWS Trainium, Azure ND H200, Groq, Together AI
  • Dedicated ASICs lead for uses needing ultra-low latency
  • Self-hosting (vLLM etc.) trade-offs vary by model and GPU price

Summary

AI competition extends to the geopolitics of silicon, power, and cooling. NVIDIA's dominance continues for now, but hyperscaler in-house chips and dedicated ASICs accelerate encroachment. The key for companies is a multi-cloud/multi-chip strategy to contain cost and supply risk.

June 2026 update

  • NVIDIA Nemotron 3 Ultra released: Now treated as the strongest U.S.-origin open reasoning model, going head-to-head with Chinese open releases like Kimi K2.6 and MiniMax M3. Chinese models still lead on some academic benchmarks, but the ceiling for U.S. open-weight reasoning models has clearly been pushed up (following the May 22 Cerebras × Kimi K2.6 throughput showdown)
  • Vera Rubin AI servers enter mass production: After the May 19 Vera CPU shipments, NVIDIA announced full mass production of the Vera Rubin platform on June 1. The Blackwell → Rubin transition is tracking faster than the original roadmap, putting long-term GPU contracts of frontier labs and hyperscalers up for review
  • RTX Spark (GB10) and GB300 Grace Blackwell Ultra expand to PCs and laptops: At Computex 2026 / GTC Taipei, NVIDIA announced an "AI agent PC" push (chasing the estimated $200B CPU market) with Microsoft, Dell, and HP. Grace Blackwell superchips are moving from cloud into clients, putting the layer where local AI agents can realistically run on Windows devices into place
  • Physical-AI and self-driving stack disclosed in one event: GTC Taipei also unveiled the "NVIDIA Cosmos 3" world model, a driving brain, an open humanoid robot platform built with Unitree, and a 32B open reasoning model targeted at robotaxis. Beyond GPUs, the "software + hardware + simulator" bundle increases NVIDIA's gravitational pull across the physical-AI stack
  • Intel announces Crescent Island GPU at Computex 2026 (up to 480GB VRAM): A single-card play to fit very large LLMs in VRAM. This is the most substantive direct counter to NVIDIA dominance in some time and gives inference-side workloads (where VRAM ceiling matters) a credible alternative to the Blackwell/Rubin path

June 2026 update

  • SpaceX leases AI compute to Google for $920M/month per SEC filing (The Decoder). The deal gives Google access to ~110,000 NVIDIA AI chips for Gemini Enterprise demand. That one of the largest cloud providers has to rent compute externally underscores how scarce AI infrastructure has become, and how tightly big-tech business lines are now intertwined. The arrangement also previews a new structural role: non-cloud GPU holders acting as wholesale GPU suppliers inside the NVIDIA economy — notable on the eve of SpaceX's IPO.
  • NVIDIA released Nemotron 3.5 ASR (MarkTechPost) — a cache-aware 600M-parameter streaming ASR model that transcribes 40 language locales in real time from a single checkpoint. Less a new research direction than a deployment-focused optimization, it signals NVIDIA pairing its GPU dominance with its own "specialized model × inference optimization" product lineup.

TechCrunch reported on 2026-08-13 that NVIDIA is preparing a new $500B plan characterized as "risky but brilliant, especially for aging GPUs" — a scheme that props up not just fresh silicon shipments but the ongoing utilization of the enormous installed base of Hopper/Ampere-class hardware. The AI-semiconductor conversation now extends beyond "this quarter's latest GPU" to the lifecycle management of millions of prior-generation accelerators already deployed worldwide.