At the Hot Chips 2026 semiconductor conference, OpenAI revealed technical specifications for its first in-house inference ASIC, codenamed Jalapeño. Developed in close partnership with Broadcom, the accelerator represents a major milestone in OpenAI's effort to curb its long-standing reliance on standard graphics processors and design bespoke datacenter silicon for high-volume model serving.
The Jalapeño processor is manufactured on TSMC's N3 process node and operates at a thermal design power of roughly 700 watts. According to OpenAI, the chip architecture focuses directly on Pareto efficiency, specifically optimizing the balance between generated tokens per joule and end-to-end inference latency under concurrent production workloads.
OpenAI stated that the custom ASIC delivers 1.5x to 1.9x efficiency gains over Nvidia GB200 and GB300 systems when executing inference tasks for ChatGPT and the Codex coding assistant. Because operational serving costs dominate day-to-day expenditures across hundreds of millions of queries, these per-token efficiency gains are intended to substantially lower operational serving expenses.
The announcement highlights a broader transition in AI compute displayed at Hot Chips 2026. While hardware companies such as Cerebras detailed roadmaps for their upcoming CS-5 and CS-6 Wafer-Scale Engines, prominent model developers are accelerating the deployment of dedicated silicon designed around their exact transformer execution requirements.
For OpenAI, Jalapeño marks an evolution toward becoming a vertically integrated compute operator. While flexible accelerator clusters remain essential for large-scale model training, regular inference deployment is migrating toward purpose-built ASICs. This development intensifies competition across the silicon market, compelling established chipmakers to prioritize real-world token economics.

