NVIDIA confirmed at GTC 2026 and Computex 2026 that the GPU architecture cadence has moved to one year, with Blackwell and Blackwell Ultra shipping in volume as of June 2026, Vera Rubin entering partner availability in H2 2026, and Rubin Ultra scheduled for H2 2027 [S1]. The top datacenter SKU shipping in Q2 2026 is the GB300 NVL72 rack, rated at 1.1 EFLOPS FP4 per rack, 130 TB/s aggregate NVLink, and 132-140 kW per rack [S1].
Blackwell Ultra B300 is the current production datacenter GPU, carrying 288GB HBM3e, 8 TB/s memory bandwidth, 15 PFLOPS dense FP4 compute, and a 1,400W TDP, with DGX B300 systems quoting 8-12 week lead times [S1]. The same B300 silicon is mirrored in the RTX PRO 6000 Blackwell workstation card (96GB GDDR7 ECC, 1.79 TB/s, 24,064 CUDA cores, 752 Tensor cores, 600W) and in the GeForce RTX 5090 consumer flagship (32GB GDDR7, 1.79 TB/s, 575W) [S1]. The H100 and H200 remain in active production for customers with deployed Hopper software stacks [S1].
Vera Rubin platform: H2 2026 production, 336B transistors, 50 PFLOPS NVFP4
Vera Rubin entered full production at GTC Taipei on June 1, 2026, with hyperscaler availability beginning H2 2026 at AWS, Google Cloud, Microsoft Azure, Oracle Cloud, CoreWeave, Lambda, Nebius, and Nscale [S1]. The Rubin GPU uses TSMC N3, carries 336B transistors, 288GB HBM4 at 22 TB/s, and is rated for 50 PFLOPS NVFP4 inference and 35 PFLOPS NVFP4 training, with NVLink 6 at 3.6 TB/s per GPU [S1].
The paired Vera CPU has 88 Olympus ARM cores (176 threads), 227B transistors, and an NVLink-C2C link at 1.8 TB/s to the Rubin GPU [S1]. A Vera Rubin NVL72 rack assembles 72 Rubin GPUs and 36 Vera CPUs into 260 TB/s aggregate NVLink, 3.6 EFLOPS NVFP4 inference, 2.5 EFLOPS training, and 100% liquid cooling [S1]. HBM4 supply is NVIDIA-certified as of June 2026 across SK Hynix, Samsung, and Micron, which is the gating item for the H2 2026 ramp [S1].
Rubin CPX and Rubin Ultra: inference specialization, then the H2 2027 step
Rubin CPX is a specialized GPU class that pairs GDDR7 memory (128GB per device) with standard Rubin GPUs in the NVL144 CPX rack, targeting the compute-bound prefill phase of million-token context inference [S1]. The CPX rack delivers 8 exaflops NVFP4 (7.5x GB300 NVL72), 100 TB per-rack memory capacity, and 1.7 PB/s per-rack memory bandwidth, with end-of-2026 availability and early AI partners Cursor, Runway, and Magic [S1].
Rubin Ultra is scheduled for H2 2027 on the same one-year cadence, with the architecture name confirmed by NVIDIA but per-SKU specifications not yet disclosed publicly as of the June 2026 update [S1]. The Feynman architecture follows in 2028, with the post-Feynman Rosa variant projected for 2029-2030 based on NVIDIA's stated annual rhythm [S1]. For context on what an annual cadence means in dollar terms, see the AI accelerator demand 2026-2030 outlook, which tracks GPU share against ASIC and HBM supply growth.
Tensor Core progression: from Volta FP16 to Blackwell 5th-gen FP8

Tensor Core capability is the single most-cited AI performance lever, and NVIDIA's published progression is: Volta (2017) introduced 1st-gen FP16/FP32 with a 12x AI training speedup vs. Pascal; Ampere (2020) added 3rd-gen structured sparsity at 2x throughput on sparse models; Hopper (2022-2024) added FP8 with dynamic range adjustment and DSMEM for cross-SM synchronisation; Blackwell (2024-2026) ships 5th-gen Tensor Cores with 2x attention-layer acceleration, 1.5x AI compute FLOPS vs. Hopper, and 25x better energy efficiency per AI operation [S2]. The Blackwell two-die GPU totals 208B transistors in its chiplet package, built around 1,422 interconnect-related patents and 617 ray-tracing patents filed 2016-2026 [S2].
That energy-efficiency curve is the underlying reason a one-year cadence is economically defensible: each step delivers enough per-operation efficiency gain to offset the rack-scale power jump from 132-140 kW (GB300 NVL72) toward 260 TB/s NVLink fabrics in Vera Rubin NVL72 [S1][S2]. The capacity story interacts directly with the DRAM 2026 HBM reallocation and 93-98% price spike, since HBM4 allocation determines how many Rubin racks ship in H2 2026.
CPU, SuperChip, and interconnect cadence: ARM, NVLink-C2C, and 1.6T switches
NVIDIA's CPU strategy is subordinated to the GPU roadmap: Grace appears inside the Grace+GPU SuperChip rather than as a standalone CPU line, evolving at a slower cadence than the GPU and pairing into GH200, GB200, and GX200 SuperChips via NVLink-C2C, then into GH200NVL, GB200NVL, and GX200NVL back-to-back modules [S3]. NVLink networks form supernodes, and larger AI clusters are stitched over InfiniBand or Ethernet fabrics [S3].
Switch silicon tracks on its own curve: Quantum-2 shipped as a 400G, 25.6T chip in 2021; Spectrum-4 enables 800G ports (51.2T) with 100G SerDes; Spectrum-5 is on the roadmap for 1.6T ports at 200G SerDes, approaching 102.4T switching capacity [S3]. NVLink and NVSwitch are expected to absorb 224G SerDes first because they are proprietary and not tied to the slower Ethernet standards cycle [S3]. SmartNIC and DPU roadmaps target ConnectX-8 and BlueField-4 at 800G, though the alignment with 1.6T switch timing is not yet public [S3].
Why a public roadmap: capacity, ceiling risk, and the competitive moat

NVIDIA publicly extends its roadmap to 2028 because customers buying several-million-dollar rack-scale systems need visible performance headroom to plan depreciation and software porting cycles [S4]. The company has noted that the Volta generation marked its AI-first pivot, and that patent analysis covering 2010-2026 identifies four discrete inflection points, anchored by the CUDA programmability moat that takes large users an estimated 6-12 months of engineering to leave [S2].
Hyperscalers are also building their own AI accelerators, so the public roadmap doubles as a commitment signal: it reminds customers that NVIDIA's system-level integration, not just the GPU die, is a multi-year investment they are buying into [S4]. The Cornell Virtual Workshop framing is useful here, since GPU architecture fundamentals explain why each generation's performance claim depends on the full pipeline (SM, HBM bandwidth, NVLink fabric, Tensor Core precision), not the transistor count alone [S5].
Decision criteria: which NVIDIA GPU tier to specify in 2026
For new datacenter purchases in Q3-Q4 2026, the three viable choices are Hopper (H100, H200, in production, existing software), Blackwell/B300 (in volume, 192-288GB HBM3e, 1,400W TDP), or waiting for Vera Rubin (H2 2026 partner availability, 50 PFLOPS NVFP4 inference, HBM4 supply-gated) [S1]. For workstation and edge inference, RTX PRO 6000 Blackwell and the RTX Spark Grace Blackwell (up to 128GB LPDDR5X unified memory, Fall 2026 launch) cover the pro and mini-PC tiers respectively [S1].
The selection criteria split cleanly along four axes: memory capacity per GPU (192GB H200, 288GB B300, 288GB Rubin, 128GB Rubin CPX GDDR7), interconnect bandwidth (130 TB/s GB300 NVL72 NVLink, 260 TB/s Vera Rubin NVL72 NVLink), precision-mode throughput (15 PFLOPS FP4 B300, 50 PFLOPS NVFP4 Rubin), and power per rack (132-140 kW GB300 NVL72, with Vera Rubin NVL72 still 100% liquid-cooled but exact kW not disclosed) [S1]. Standard NVIDIA guidance: specify the HBM generation and supplier list, pin the rack kW budget, and confirm the NVLink topology before locking a multi-year capacity plan, since HBM4 supply and Rubin Ultra specs are the two largest unknowns on the 18-month horizon [S1][S3].
Trackable signals through Q1 2027: NVIDIA's first Rubin Ultra SKU disclosure, HBM4 volume yield reports from SK Hynix, Samsung, and Micron, and the first public Vera Rubin NVL72 customer shipment announcements from AWS, Google Cloud, or Microsoft Azure [S1]. The server hardware 2026 hyperscaler capex map is the most relevant cross-reference, since GPU rack orders and DRAM allocation move together inside the same hyperscaler procurement cycle.
Spec-level background on the components involved: pressure transmitter, flow meter, and industrial valve.