← Back to feed Article · August 26, 2026 · 3 min
Articles

Silicon Rivalry: OpenAI's Jalapeño ASIC Outpaces Nvidia Blackwell in Efficiency

OpenAI and Broadcom have developed a dedicated inference processor called Jalapeño that beats Nvidia's flagship Blackwell hardware in tokens per megawatt. The custom silicon aims to slash data-center power costs and reduce generation latency across AI services.

Photo: Хабр ML

Sixteen months — that is all it took for OpenAI from hiring its first hardware engineers in mid-2024 to getting finished silicon in hand. Aiming to break free from Nvidia's exorbitantly priced server accelerators, the company developed its own custom chip, dubbed Jalapeño, in partnership with Broadcom. Unveiled at the Hot Chips conference, the hardware was tested hands-on by SemiAnalysis analysts in the developers' lab using the InferenceX benchmark suite. The biggest surprise: the ChatGPT creator's very first silicon outperformed not only offerings from AMD and Google, but also Nvidia's flagship Blackwell architecture in token throughput per megawatt.

Why build custom silicon

The rationale is strictly practical. Standard Nvidia GPUs are general-purpose workhorses designed to handle everything from 3D gaming graphics to training massive neural networks from scratch. Jalapeño was architected as an application-specific integrated circuit (ASIC) solely for inference — generating finished answers for end users. According to SemiAnalysis, the chip incorporates high-speed HBM4 memory and avoids proprietary lock-in: it comfortably ran third-party open-weight models including DeepSeek R1, Kimi-K2.5, and GPT-OSS. To demonstrate its versatility, engineers even used Codex prompts to get the classic game Doom running on the silicon.

Photo: Habr ML

In hands-on testing, the performance gains are striking. Running DeepSeek R1 in single-user mode, Jalapeño delivers over 700 tokens per second, climbing to roughly 1,400 tokens per second per user on Kimi-K2.5 and GPT-OSS. Notably, the chip sustains these speeds on standard baseline algorithms without software shortcuts, speculative decoding, or complex compute-phase splitting.

Jalapeño outperforms competing chips in perf/W, delivering superior token throughput per megawatt across the entire system.

Meanwhile, GSM8k benchmark evaluations confirmed that Jalapeño matches Nvidia's flagship silicon in output precision and reasoning quality.

What changes on user screens

Translating engineering specs into everyday terms, this silicon rivalry yields two clear benefits for anyone using AI tools daily. First, it eliminates irritating latency. During peak hours, when millions of users simultaneously generate text or debug code, servers powered by dedicated ASICs avoid bottleneck slowdowns. That familiar lag where an answer painstakingly renders character by character stems directly from server bandwidth limits.

The second benefit is cost. Running massive data centers packed with Nvidia accelerators costs developers billions, and that monopoly premium ultimately gets passed down to consumers through rising subscription fees. An energy-efficient processor that consumes less electricity per generated word directly lowers infrastructure operating costs. SemiAnalysis analysts note one caveat: current figures come from OpenAI's test benches, and independent benchmarks handling long context windows are still pending.

Photo: Habr ML

OpenAI has demonstrated that Nvidia's market dominance is vulnerable. If expensive green-team hardware loses its status as the default data-center standard, daily interactions with ChatGPT will become noticeably snappier while subscription prices stabilize.

Source Habr ML → © 2026 «Gadgety». Full or partial copying — with a link to this page.