OpenAI’s Jalapeño AI chip packs 13.4 petaFLOPS

Home Semiconductor News OpenAI’s Jalapeño AI chip packs 13.4 petaFLOPS
OpenAI's new inference chip, Jalapeno

OpenAI’s Jalapeño Intelligence Processor targets cheaper token serving and Nvidia independence

In sum — what we know

  • An aggressive design cycle — OpenAI moved Jalapeño from design to manufacturing tape-out in about nine months, claiming the chip was designed in part by AI
  • The performance claims — A 700W Jalapeño reportedly delivers 1.5x–1.9x higher performance-per-watt than Nvidia’s GB200 and GB300 on SemiAnalysis’s InferenceX suite
  • The deployment timeline — Initial rollout lands in late 2026 at small volumes, scaling through 2027 for production workloads

OpenAI is offering some more details about its “Jalapeño” custom AI chip, or what it calls an “Intelligence Processor.” At Hot Chips 2026, the company detailed specifications, its timeline for deployment, and more. According to OpenAI, Jalapeño went from initial design to manufacturing tape-out in roughly nine months, an unusually rapid schedule for a high-end ASIC, and OpenAI leaned into that with a claim that the chip was designed (in part) by AI. 

The architecture

Jalapeño is a reticle-sized design, meaning it occupies roughly the maximum area that can be printed in a single photolithography exposure. That’s about as aggressive as monolithic chip design gets. Each accelerator delivers 13.4 petaFLOPS of mixed-precision matrix compute using mxfp4 x mxfp4 — a 4-bit floating-point format tuned specifically for LLM workloads. The narrowness is the point. By stripping out training flexibility and everything else Nvidia has to support in a general-purpose GPU, OpenAI could spend the entire die on dense matrix units and memory built for serving tokens.

The memory system pairs multiple stacks of HBM4 for 216 GB of capacity and 15.4 TB/s of bandwidth. All of this runs at roughly 700 watts, which sounds like a lot until you compare it to Nvidia’s flagship GPUs operating in the 1,200W to 1,400W range. That power gap does a lot of heavy lifting in OpenAI’s efficiency claims, as we’ll get to.

System scaling

The chip is designed strictly for large-scale rack deployment, and the base unit is a 128-accelerator system delivering 1.7 exaFLOPS of 4-bit compute, 27.5 TB of HBM4, and nearly 2 petabytes per second of aggregate memory bandwidth. From there, the architecture scales to a 2,048-chip domain aimed at frontier models — roughly 27 exaFLOPS of compute and 432 TiB of memory across the full system.

Holding this together is a two-tier interconnect. Each accelerator gets 600 GB/s of bandwidth for local communication within its 128-chip pod, and 200 GB/s for global communication across the full 2,048-chip domain. The logic is straightforward. The chips jointly serving a single request need low latency and high bandwidth between them, while the broader fabric only needs enough throughput to scale out across racks. It’s a pragmatic design, and it reflects the reality that frontier models now run across clusters, not devices.

Big performance

The numbers OpenAI presented are striking, if you take them at face value. The company says Jalapeño lowers latency by 1.7x compared to an Nvidia GB200, while increasing tokens-per-second by 1.9x per kilowatt of power. On SemiAnalysis’s public “InferenceX” benchmark suite, a 700W Jalapeño reportedly delivers 1.5×–1.9× higher performance-per-watt than Nvidia’s GB200 and GB300 accelerators, with end-to-end latency 1.7×–3.6× lower on comparable workloads. For an interactive product like ChatGPT, that latency figure arguably matters more than raw throughput.

But some caution is warranted. These are vendor-run benchmarks on early silicon, plus one specific public test suite — a combination that may not reflect the heterogeneous mess of real-world data center workloads. And the comparisons pit a 700W part against Nvidia hardware drawing up to twice the power, which naturally flatters Jalapeño’s per-watt metrics while leaving absolute performance questions open. 

Deployment plans

Initial deployment is slated for the end of 2026 in very small volumes, mostly for validation and limited production workloads, with the real rollout expanding through 2027 and beyond. That measured pace makes sense given the compressed design cycle. Longer term, Jalapeño is expected to power OpenAI’s next-generation, gigawatt-scale “Stargate” data centers being built with Microsoft, Oracle, and SoftBank.

Developers will never touch this hardware directly, but they may feel it. If the cost-per-token reduction holds up, it could eventually translate into cheaper API pricing. Of course, more likely is that pricing will stay as it is, allowing OpenAI to pocket the savings.

Jalapeño also fits a well-established pattern. Google has its TPUs, Amazon has Trainium and Inferentia, Meta has MTIA — every hyperscaler with the resources to build custom silicon is doing so to reduce its dependence on Nvidia. OpenAI is now firmly in that club, though only halfway. The company remains entirely dependent on Nvidia and other accelerators for training, and that split creates its own friction. Models trained on one architecture and served on another can complicate optimization workflows in ways that don’t show up on a benchmark slide. Jalapeño is a real step toward hardware independence, but it’s a first step.

What you need to know in 5 minutes

Join 37,000+ professionals receiving the AI Infrastructure Daily Newsletter

This field is for validation purposes and should be left unchanged.

This website uses cookies to improve your experience. We'll assume you're ok with this, but you can opt-out if you wish. Accept Read More