Skip to content
AI ConnectPowered by VELENTIS
AI-generated1 min

Hot Chips 2026: OpenAI Unveils Custom Jalapeño Inference ASIC

At Hot Chips 2026, OpenAI revealed details of its custom Jalapeño inference ASIC developed with Broadcom, promising notable efficiency gains over Nvidia GB200 and GB300 systems.

This article was AI-generated and published automatically. Context, labelling and all sources at the end of the article.

(KI-generiertes Symbolbild: Gemini / AI Connect)

At the Hot Chips 2026 semiconductor conference, OpenAI revealed technical specifications for its first in-house inference ASIC, codenamed Jalapeño. Developed in close partnership with Broadcom, the accelerator represents a major milestone in OpenAI's effort to curb its long-standing reliance on standard graphics processors and design bespoke datacenter silicon for high-volume model serving.

The Jalapeño processor is manufactured on TSMC's N3 process node and operates at a thermal design power of roughly 700 watts. According to OpenAI, the chip architecture focuses directly on Pareto efficiency, specifically optimizing the balance between generated tokens per joule and end-to-end inference latency under concurrent production workloads.

OpenAI stated that the custom ASIC delivers 1.5x to 1.9x efficiency gains over Nvidia GB200 and GB300 systems when executing inference tasks for ChatGPT and the Codex coding assistant. Because operational serving costs dominate day-to-day expenditures across hundreds of millions of queries, these per-token efficiency gains are intended to substantially lower operational serving expenses.

The announcement highlights a broader transition in AI compute displayed at Hot Chips 2026. While hardware companies such as Cerebras detailed roadmaps for their upcoming CS-5 and CS-6 Wafer-Scale Engines, prominent model developers are accelerating the deployment of dedicated silicon designed around their exact transformer execution requirements.

For OpenAI, Jalapeño marks an evolution toward becoming a vertically integrated compute operator. While flexible accelerator clusters remain essential for large-scale model training, regular inference deployment is migrating toward purpose-built ASICs. This development intensifies competition across the silicon market, compelling established chipmakers to prioritize real-world token economics.

What this means for you

For enterprise developers and end users, custom silicon like Jalapeño points toward lower API latency and more resilient pricing structures for high-volume model calls. The transition underscores that scaling frontier AI sustainably requires hardware co-design rather than relying solely on general-purpose GPUs. This shifts long-term bargaining power within the AI supply chain.

Perspectives

Coverage: 4× Other

One story, several angles: how each source frames the topic, each with a verbatim quote.

  • tomshardware.comOther

    Tom's Hardware highlights that the AI-developed Jalapeno ASIC achieves notable efficiency and throughput gains against Nvidia's power-hungry Blackwell processors.

    Original quote

    accelerator developed using AI achieves efficiency and throughput gains against power-hungry Blackwell

    tomshardware.com
  • servethehome.comOther

    ServeTheHome focuses on the detailed technical architecture and benchmarks of the Jalapeno system co-developed with Broadcom, framing it as a complete inference platform optimized for performance per watt and low latency.

    Original quote

    OpenAI Jalapeño is framed as an inference platform rather than a raw accelerator.

    servethehome.com
  • latent.spaceOther

    Latent Space frames Jalapeno as a major Blackwell-beating alternative, emphasizing its efficiency metrics alongside the use of AI models to write and optimize low-level kernels.

    Original quote

    OpenAI released first benchmark details for its custom inference chip Jalapeño

    latent.space

Source classification is maintained editorially (political spectrum only where consensus is broad; vendor communication is PR, not journalism). Unlabelled sources are unclassified: we do not guess.

Evidence

Well sourced
73/100

The evidence score is computed, not hand-set: from confidence, the number of sources and the share of verified statements.

Source & transparency

As of: August 27, 2026

AI-generatedAI-generated: produced automatically from vetted sources with technical quality checks (source, quote and figure verification); no human sign-off of each item before publication

Sources
4
Verified statements
2 / 4
Evidence score
73Well sourced

Want to put this into practice?

We connect you with suitable AI providers from the DACH region, free of charge and without obligation.

What's next?