OpenAI has published initial performance results for its custom Jalapeño AI inference chip. The benchmarks, presented at the Hot Chips conference, indicate that Jalapeño delivers faster responses and higher throughput for AI inference tasks.
Richard Ho, OpenAI's head of hardware, stated that Jalapeño offers both lower latency and higher throughput, a combination that AI systems typically require a trade-off between. The chip was developed in collaboration with Broadcom and was first announced in June.
Jalapeño was tested using the InferenceX benchmarking platform, comparing its performance against Nvidia's GB200 or GB300 superchips. OpenAI reported that Jalapeño achieved 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower latency.
The chip also demonstrated more tokens per user and throughput per kilowatt than current state-of-the-art inference processors. These results were observed across models such as GPT-OSS 120B, DeepSeek R1, and Kimi K2.5.
The development of Jalapeño specifically targets the efficiency of large language model operations, particularly for multi-step agent tasks. AI agents often call models repeatedly, leading to compounding delays over the course of a longer task.
Jalapeño's design aims to provide higher throughput without increasing response times, which is critical for agents that need to complete many steps in sequence.
OpenAI anticipates that Jalapeño will see initial deployment in small volumes by the end of 2026, with more significant deployment expected in 2027. The company's own AI models assisted in the development process of the chip.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
OpenAI presented benchmark results for its Jalapeño chip at the Hot Chips conference, demonstrating improved tokens per user and throughput per kilowatt compared to current state-of-the-art inference processors. This development indicates progress in specialized hardware for AI inference, potentially leading to more efficient and faster AI model deployment.
OpenAI published initial performance results for its custom Jalapeño inference chip, developed with Broadcom, across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 models. The results indicate higher throughput without increased response times, specifically addressing the compounding delays experienced by AI agents. This development is significant for improving the efficiency of large language model operations, particularly for multi-step agent tasks.
OpenAI announced its new Jalapeño AI chip, developed in partnership with Broadcom, delivers faster responses and higher throughput for AI inference tasks compared to Nvidia's superchips. The chip showed 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower latency in benchmark tests. This development could lead to more responsive AI systems and improved access as demand for AI grows.