OpenAI’s Jalapeño ASIC delivered 1.7‑3.6 times lower end‑to‑end latency than Nvidia’s GB200 and GB300 superchips in benchmark tests released on Aug. 25, 2026, according to a Verge report that quoted OpenAI’s own blog and a briefing with hardware vice‑president Richard Ho.
Background on the players
OpenAI, the San Francisco‑based artificial‑intelligence firm founded on Dec. 11, 2015, is led by chief executive Sam Altman and employs roughly 4,500 people (Wikidata). The company has been expanding beyond software into custom silicon to accelerate its own models. In June 2026 the firm announced its first in‑house ASIC, named Jalapeño, in a blog post that highlighted a partnership with Broadcom to fabricate the chip.
Nvidia Corp., headquartered in Santa Clara, is the dominant supplier of AI‑focused GPUs and, more recently, purpose‑built inference superchips. Jensen Huang remains its chief executive. In its most recent 10‑Q filed Aug. 26, 2026 the company reported $177.8 billion in revenue for the fiscal year ending July 26, 2026 and $118.0 billion in net income, underscoring the scale of the market it serves.
Benchmark methodology and results
The performance numbers come from the InferenceX platform, the industry‑standard inference benchmark suite cited in the packet’s supporting claims. OpenAI ran Jalapeño against the best recorded results at the time – Nvidia’s GB200 and GB300 superchips – across three high‑profile large‑language‑model workloads: GPT‑OSS 120B, DeepSeek R1 and Kimi K2.5 1T.
The Verge article (2026‑08‑25) reproduced OpenAI’s statement: “Jalapeño delivered 1.5 to 1.9 times more AI work per watt across GPT‑OSS 120B, DeepSeek R1, and Kimi K2.5 1T than the comparison systems, while offering 1.7 to 3.6 times lower end‑to‑end latency across the three models.” The same piece quoted Richard Ho, OpenAI’s hardware vice‑president, saying Jalapeño offers the “best of both worlds” – lower latency without sacrificing throughput.
In concrete terms, the latency advantage means that for a given inference request the Jalapeño‑powered system returns a response up to 3.6 times faster than a comparable Nvidia‑based system. The work‑per‑watt metric indicates that the chip can perform up to 1.9 times more AI work for each watt of power consumed, a key efficiency lever for data‑center operators.
Competitive context
Nvidia’s GB200 and GB300 superchips have been the reference point for high‑throughput inference since their launch earlier in 2026. Their performance figures were publicly disclosed in Nvidia’s product briefings and have been used as the baseline in most third‑party benchmark suites. By positioning Jalapeño against these specific chips, OpenAI is directly challenging Nvidia’s claim to be the sole provider of world‑class inference hardware.
The packet notes that the benchmarks were run on the InferenceX suite, which standardises model inputs, batch sizes and measurement methodology. This eliminates the “cherry‑picking” criticism that sometimes clouds hardware comparisons, because the same workload and measurement window were applied to both Jalapeño and the Nvidia reference chips.
Implications for the AI hardware market
If the reported latency and efficiency gains hold in production, data‑center operators could see two immediate benefits. First, lower latency translates into faster user‑facing responses for applications such as chatbots and real‑time translation. Second, higher work‑per‑watt reduces electricity costs and eases thermal constraints, allowing more dense deployments.
OpenAI plans to ship Jalapeño in limited volumes by the end of 2026 and to ramp up production in 2027, according to the company’s timeline. The chip’s partnership with Broadcom – a major semiconductor foundry – suggests that OpenAI is leveraging established manufacturing capacity rather than building a fab from scratch.
For Nvidia, the results raise the question of how quickly it can iterate on its own inference silicon. The company’s recent 10‑Q filing shows a massive cash flow that could fund accelerated R&D, but the packet does not contain any forward‑looking statements from Nvidia about a response to Jalapeño.
What remains unknown
- Pricing: Neither OpenAI nor Broadcom disclosed the cost per chip, so the economic trade‑off for customers is unclear.
- Scale of production: The “small volumes” slated for late‑2026 have no disclosed unit count.
- Software stack integration: While the benchmark used the InferenceX suite, the packet does not detail how Jalapeño will integrate with existing OpenAI model serving frameworks.
- Long‑term reliability: No data on yield rates or failure modes were provided.
OpenAI’s blog post, referenced by the Verge, confirms the latency and efficiency numbers but does not elaborate on these open questions. Future updates from the company’s engineering team will be needed to assess the chip’s commercial viability.
Benchmark performance table
| Model | Latency (× lower vs Nvidia) | Work‑per‑Watt (× higher vs Nvidia) |
|---|---|---|
| GPT‑OSS 120B | 1.7‑3.6 | 1.5‑1.9 |
| DeepSeek R1 | 1.7‑3.6 | 1.5‑1.9 |
| Kimi K2.5 1T | 1.7‑3.6 | 1.5‑1.9 |
Source: The Verge – OpenAI Jalapeño chip benchmarks (2026‑08‑25).
Looking ahead
OpenAI’s Jalapeño ASIC marks the first publicly disclosed silicon that claims a multi‑fold latency advantage over Nvidia’s current flagship inference chips. The next data points to watch are the first customer deployments slated for Q4 2026 and any performance updates from Nvidia’s upcoming GPU generations. As the AI inference market matures, hardware efficiency will become a decisive factor for cloud providers, and Jalapeño’s early results suggest a new competitive axis beyond raw throughput.
Stakeholders – from data‑center operators to venture investors – should monitor OpenAI’s production schedule and any pricing signals that emerge in the coming months. Until then, the benchmark numbers remain the strongest publicly available evidence of Jalapeño’s advantage.