OpenAI’s Jalapeño AI Chip

OpenAI has released the first performance results for Jalapeño, its custom artificial intelligence inference chip developed with Broadcom, reporting gains in processing speed and power efficiency against commercially available AI hardware.

The company tested Jalapeño using SemiAnalysis’ public InferenceX benchmark across three models: GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. According to OpenAI, the chip delivered between 1.5 and 1.9 times higher peak performance per watt and between 1.7 and 3.6 times lower end-to-end latency than the comparison systems across the three models.

For highly interactive workloads, Jalapeño recorded between 2.1 and 4.1 times higher performance, according to the company’s testing. On Kimi K2.5 1T, the largest public model evaluated, OpenAI reported approximately 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than the comparison system.

Jalapeño has a rated power consumption of 700 watts, although OpenAI said sustained power remained at or below 550 watts during the workloads tested. The company said the architecture is designed to reduce data movement and communication delays, allowing compute, memory and networking resources to be used differently depending on the stage of inference.

The processor was co-developed with Broadcom and forms part of OpenAI’s broader effort to build more of the infrastructure supporting its AI models and products. Jalapeño was developed from initial design to manufacturing tape-out in nine months, with OpenAI’s own AI models used during parts of the chip design and optimisation process.

The company also used Codex with GPT-Astra to optimise software for the processor. OpenAI said three open-weight models were brought to high performance within two months, while AI-generated implementations for selected GPT-OSS attention and mixture-of-experts blocks ran between 1.5 and 1.8 times faster than existing human-written implementations.

OpenAI plans to begin deploying Jalapeño within its infrastructure in limited volumes by the end of 2026, with larger-scale deployment expected in 2027. The company is already developing a second generation, while plans for a third generation are also taking shape.

The custom processor, however, is not expected to replace OpenAI’s existing chip suppliers entirely. OpenAI said it will continue deploying accelerators from Nvidia and other partners for training and inference alongside its own hardware.

The benchmark results mark the first measurable indication of how OpenAI’s move into custom silicon could influence the infrastructure behind its AI services, as AI companies increasingly focus on reducing inference costs while improving response speeds and capacity.

Disclaimer: This article may include information derived from interviews, press releases, public statements, research, company communications and other publicly available or third-party sources. Such material may be summarised, paraphrased or contextualised for journalistic and editorial purposes. All rights in third-party content remain with their respective owners.