Nvidia reportedly revives Rubin CPX with HBM4 redesign

Home Semiconductor News Nvidia reportedly revives Rubin CPX with HBM4 redesign
Nvidia Vera Rubin Feynman Rubin CPX

Analyst reports 168 GB of HBM4, standalone racks, and 2027 production

In sum – what we know:

  • Supply-chain revival — Ming-Chi Kuo reports Rubin CPX is back on a production schedule beginning in Q1 2027, after Nvidia dropped it from the GTC 2026 roadmap in favor of Groq 3 LPX racks. (unchanged)
  • A major redesign — According to Kuo, the chip abandons the shelved design’s 128 GB of GDDR7 for 168 GB of HBM4, a 2,300 W power envelope matching Rubin, and its own standalone MGX ETL racks.
  • Split inference economics — Per the supply-chain reporting, CPX handles the compute-bound prefill stage while Rubin GPUs decode over Ethernet RDMA, at Nvidia’s recommended 1:1 CPX-to-GPU pairing.

Nvidia’s Rubin CPX accelerator is apparently back from the dead. According to supply-chain analyst Ming-Chi Kuo, the specialized AI chip is active again and entering a new production schedule, after much of the market assumed it was dead — a plausible assumption, given that Nvidia dropped Rubin CPX from its GTC 2026 keynote and roadmap in March. Kuo says his latest industry checks point to production beginning in the first quarter of 2027, which implies commercial deployment later that year.

Kuo says the chip has undergone a substantially redesigned architecture compared with the original plan. Detailed specifications of the redesign have surfaced in Kuo’s reporting — a move to HBM4 memory, a new standalone rack design, and a higher power envelope — but none of it is officially confirmed by Nvidia, so it all rests on his supply-chain checks. That’s worth keeping in mind, because the Q1 2027 production target represents a slip of at least several months from Nvidia’s own earlier messaging, which had Rubin CPX available around the end of 2026.

Performance specs

Rubin CPX is a purpose-built accelerator aimed at the compute-bound “prefill” stage of inference, where a model ingests an enormous prompt before it starts generating tokens. As context windows stretch toward a million tokens, that prefill phase starts to dominate the compute cost of a query, which is the argument for dedicating specialized silicon to it.

According to Kuo, the redesigned Rubin CPX abandons the GDDR7 that the original 2025 design carried (roughly 128 GB per chip) and moves to 168 GB of HBM4 per chip — less than the 288 GB on the main Rubin GPU, but far more bandwidth than the shelved design offered — while its power envelope rises to 2,300 W, matching the standard Rubin GPU’s ceiling. The memory shift makes the cost story more complicated: HBM4 is significantly more expensive than GDDR7, so the case for CPX now rests on prefill efficiency rather than on cheaper memory. Each chip in the original announcement was rated at around 30 petaFLOPS in Nvidia’s NVFP4 format, with Nvidia claiming roughly three times the attention-mechanism throughput of a Blackwell GB300 GPU in prefill-centric workloads — but those figures date to September 2025 and are likely superseded by the redesign; Kuo’s reporting offers only that the revived CPX delivers near-Rubin compute performance and stronger prefill throughput, with no new official numbers yet.

The architectural role is arguably the more interesting part. CPX handles prefill while the main Rubin GPU takes over token-by-token decoding — the first time Nvidia has split these two phases across distinct chip designs. In theory, that keeps expensive HBM-based GPUs from burning cycles on a front-loaded compute spike they’re overbuilt for.

We don’t yet have the full details on why Nvidia is pursuing the redesign in the first place. Coverage points to a few plausible candidates — like cost pressures around GDDR7 supply, evolving customer requirements for million-token models, or bottlenecks that surfaced during early Rubin system testing that didn’t match original expectations. Any of these would justify reworking die partitioning, memory configuration, or process node choices, The design details that have emerged — the HBM4 memory, the 2,300 W power envelope, the standalone rack architecture — all come from Kuo’s supply-chain conversations rather than anything on an official spec sheet, and any of the drivers above could have motivated the changes they describe.

Platform integration

CPX is a companion chip within the broader Vera Rubin platform, sitting alongside the main Rubin GPU, the Vera CPU, CX9 NICs, and Spectrum-X switches. The intended deployment was originally co-located: Nvidia’s 2025 announcement described Vera Rubin NVL144 CPX racks, where CPX cards sat alongside standard Rubin GPUs interconnected via NVLink and Spectrum-X. The redesign moves CPX into its own standalone MGX ETL racks instead — configurations of 64, 128, 192, or 256 CPX GPUs, built from 64-GPU modules that each hold eight 8-GPU compute trays plus a switch tray. Within each tray, the CPX GPUs connect via NVLink; across the rack, scale-out runs over Spectrum-6 Ethernet, with all-copper links inside a module and optical OSFP links between modules. The CPX racks then hand off context to Vera Rubin NVL72 decode racks over Ethernet RDMA, with Nvidia recommending a 1:1 ratio of CPX GPUs to Rubin GPUs for the pairing.

Splitting prefill and decode across two chip designs means Nvidia has to adapt its software stack to schedule execution across both parts. That’s more complex than existing multi-GPU strategies, and framework and model developers will likely need to optimize for the split themselves to see the full benefit, particularly for custom attention mechanisms and retrieval-heavy pipelines. A redesigned chip arriving in 2027 gives that software work time, but it also adds validation risk on top of an already slipped schedule.

The competitive backdrop makes CPX strategically important regardless. By 2027, Nvidia will be contending with AMD accelerators optimized for large-context and memory-bound workloads, so delivering a genuine TCO advantage with CPX matters for holding onto its inference leadership. Hyperscalers have been pushing for exactly this kind of specialized, cost-optimized accelerator for very long-context inference, and they’re the obvious early adopters. Enterprise users will probably see CPX-based cloud instances sometime after that, depending on how quickly the redesigned part gets integrated.

Whether CPX becomes central to Nvidia’s portfolio is still an open question, though. Hyperscalers could embrace the specialized prefill workflow, or they could decide the added complexity isn’t worth it and simply scale out more standard Rubin GPUs.

What you need to know in 5 minutes

Join 37,000+ professionals receiving the AI Infrastructure Daily Newsletter

This field is for validation purposes and should be left unchanged.

This website uses cookies to improve your experience. We'll assume you're ok with this, but you can opt-out if you wish. Accept Read More