OpenAI Benchmarks Jalapeño Against Nvidia AI Chips
OpenAI says its first inference chip delivers lower latency and more work per watt across three large open-weight models.
Measured silicon, not another preview
OpenAI has published the first detailed performance results for Jalapeño, its custom inference accelerator developed with Broadcom. On the public InferenceX benchmark, the company tested complete serving systems running GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, spanning both high-throughput and latency-sensitive operating points.
OpenAI reports that Jalapeño completed 1.5 to 1.9 times more work per watt at peak throughput than systems based on Nvidia GB200 or GB300 accelerators. End-to-end latency was between 1.7 and 3.6 times lower, while performance in highly interactive configurations improved by 2.1 to 4.1 times. On Kimi K2.5, for example, OpenAI measured about 1.5 times higher peak performance per watt and 3.4 times lower latency than the comparison system.
The company normalized results using published package power ratings: 700 watts for Jalapeño, 1,200 watts for GB200 and 1,400 watts for GB300. It said Jalapeño sustained no more than 550 watts during the measured workloads. These remain vendor-produced results, however, and production-scale reliability, software maturity and total system cost have not yet been independently demonstrated.
A bid to control inference economics
The design keeps model state and KV cache close to the relevant compute resources while coordinating memory, networking and processing for both prompt ingestion and token generation. OpenAI also said AI-assisted engineering helped take the chip from initial design to tapeout in nine months; selected AI-generated kernels ran 1.5 to 1.8 times faster than expert-written versions, though that comparison did not cover entire models.
OpenAI plans to deploy Jalapeño internally by year-end while continuing to buy Nvidia and other accelerators. The significance is strategic: credible first-party inference silicon could reduce OpenAI's marginal serving costs, improve agent responsiveness and give it greater leverage over the hardware suppliers underpinning its products. The decisive evidence will come after deployment, when utilization, yield and operating costs can be observed outside a controlled benchmark.