What happened?

On 25 August 2026, at Hot Chips, OpenAI published the first measured numbers for Jalapeño, its first custom inference ASIC, co-developed with Broadcom and racked with Celestica. The June 24 unveil had been a photo and a promise. This week is silicon in the lab, running other people's models, on a public serving bench.

The chip is not a GPU. It is a blank-sheet inference processor: 216 GiB of HBM4, 15.4 TB/s of bandwidth, 13.4 PFLOP/s of mxfp4 matrix compute, 700 W package TDP. OpenAI says sustained draw on the tested workloads stayed at or below 550 W. A local domain is 128 ASICs at 600 GB/s; a global domain of 2,048 chips rides Broadcom Tomahawk6 Ethernet at 200 GB/s. That is a TPU-shaped scale-up story, not an NVL72 clone.

Tape-out was claimed at nine months, with OpenAI models in the design loop (XLS, arithmetic circuits, later Codex + GPT-Astra for kernels). Richard Ho, Ravi Narayanaswami, and Chris Leary walked the slides. Architecture concept late 2024, RTL freeze 2025, A0 silicon then ChatGPT on the die. Gen 2 is "deep in development." Gen 3 is "taking shape." Ho's deployment calendar is the part that should survive the charts: very small volumes at end of 2026, a real ramp in 2027. No merchant SKU. Internal demand eats the wafers.

The commercial envelope around the die is older and larger. On 13 October 2025, OpenAI and Broadcom announced a collaboration for 10 gigawatts of OpenAI-designed accelerators, racks targeted to start H2 2026 and complete by end of 2029. Jalapeño is Gen 1 of that bet, not a side project.

The numbers OpenAI actually published

Bench is InferenceX (SemiAnalysis), nominal 8k input / 1k output, single-token prediction (STP), results normalized to published package TDP. Comparison systems: GB200 at 1,200 W (GPT-OSS 120B) and GB300 at 1,400 W (DeepSeek R1 670B MXFP4, Kimi K2.5 1T MXFP4). Headline across the three models: 1.5–1.9× more mixed tokens per kilowatt at peak, 1.7–3.6× lower end-to-end latency, 2.1–4.1× on the interactive (min time-between-tokens) end.

Matched points OpenAI printed:

Model System Peak mixed TPS / kW End-to-end latency Min TBT Interactivity (tok/s/user)
GPT-OSS 120B Jalapeño 85,448 1.03 s 0.69 ms 1,459
GB200 44,960 1.80 s 1.87 ms 535
Gain ~1.9× ~1.7× ~2.7× ~2.7×
DeepSeek R1 670B Jalapeño 19,641 1.65 s 1.43 ms 700
GB300 11,781 5.99 s 5.90 ms 169
Gain ~1.7× ~3.6× ~4.1× ~4.1×
Kimi K2.5 1T Jalapeño 18,195 1.56 s 1.44 ms 694
GB300 11,862 5.31 s 5.48 ms 182
Gain ~1.5× ~3.4× ~3.8× ~3.8×

OpenAI also printed a 50–100× column: throughput at the previous best time-between-tokens, ~53.7× for GPT-OSS 120B on GB200, ~104.3× for DeepSeek R1 on GB300, ~56.1× for Kimi K2.5. That is not "the chip is 100 times faster." It is throughput at the other system's best interactivity. Blackwell can go fast, or it can go dense. Jalapeño's claim is it does not have to pick. Ho was explicit that STP vs multi-token prediction is the other distortion: NVIDIA public numbers often include MTP (he put that boost at 3–5×). OpenAI also ran a DeepSeek comparison of Jalapeño STP against GB300 MTP and still claimed the frontier. Treat that as first-party, lab silicon, one sequence length.

Codex + GPT-Astra brought the three open-weight models up in about two months after A0. Selected GPT-OSS attention and MoE blocks, AI-written, ran 1.5–1.8× faster than the human-expert kernels. That is blocks, not the full model. It is still the interesting software fact: the programming model (local tensors, explicit communication, Gluon spatial cores) was built so frontier models can write the kernels.

Why this is not "game over for Nvidia"

The thumbnail is a funeral. The Hot Chips talk is not.

  • Wrong generation, wrong job. SemiAnalysis, after sitting in the lab, told CNBC the Blackwell comparison is "somewhat incomplete and unfair": Jalapeño has HBM4, Blackwell does not. Vera Rubin is the like-for-like, and Rubin systems are shipping to customers now, while Jalapeño is still engineering samples. Inference is the slice. Training stays on GPUs. Ho said NVIDIA (and others) remain in the fleet.
  • One published recipe. OpenAI showed 8k/1k. InferenceX also has AgentX (long-context, multi-turn, tool-call shaped traffic). That is the workload agents actually generate. Until those numbers are as public as the 8k/1k charts, the "responsive agents" line in the blog is a thesis, not a measurement.
  • Captive silicon. Ho: external sales are not the point; OpenAI's own demand will consume capacity. Same pattern as Maia, MTIA, TPU. You cannot rent Jalapeño. The unit-economics win, if it holds at rack scale, shows up as OpenAI opex and API price, not as a new instance type.
  • Same bottleneck queue. TrendForce, citing Tom's Hardware, puts Gen 1 on TSMC N3 with six HBM4 stacks. Samsung is the rumored HBM supplier. 10 GW of this die is another claimant on the same wafers, HBM, and CoWoS that NVIDIA already books. Custom architecture does not create extra EUV tools.
  • ASICs bet the model. Prefill is compute-bound, decode is bandwidth-bound, MoE is bursty collectives. OpenAI's answer is one balanced chip that gates idle blocks and keeps KV local, instead of a heterogeneous fleet whose idle accelerators still burn HBM and network power. That is a good 2026 bet. It is a bad 2028 bet if the model layer moves in a way the die did not provision for. CUDA's moat is not peak FLOPS. It is that the next architecture still compiles.

Omdia told CNBC it expects custom ASICs to exceed GPUs in volume by 2028, with GPU revenue lagging because GPUs stay expensive. Yole called Jalapeño a threat to NVIDIA inference margins, not to CUDA training lock-in.

Why this is interesting

Inference is the bill that never stops. Training is a project. ChatGPT, Codex, and API agents are a factory. OpenAI's own metrics for the factory are time to last token and tokens per joule, not TFLOPS on a slide. Agents make that worse: each tool round trips the latency. A chip that is merely "fast at batch" is a cost center for that product.

The second-order story is the loop. Models helped design the chip; the chip is being programmed by models; the serving stack is being co-designed with the product. That is what "full stack" means when it is not a slogan. It is also why 10 GW through 2029 matters more than any one InferenceX point. If tokens per watt move even 20–30% at that power envelope, the Stargate financing math in the datacenter war note changes. If they do not, Jalapeño is an expensive way to learn that Broadcom + TSMC + HBM4 was always the scarce input, which is the same lesson Terafab is pouring concrete about from the other side of the foundry.

For everyone building agents on APIs: you do not get this silicon. You get whatever price and latency OpenAI can extract once a few racks are qualified. Watch 2027 volumes, AgentX (or equivalent long-running session numbers), and whether API $ / token actually moves. The die is real. The funeral is content.