OpenAI has published initial performance results for its custom Jalapeño AI inference chip. The benchmarks, presented at the Hot Chips conference, indicate that Jalapeño delivers faster responses and higher throughput for AI inference tasks.
Richard Ho, OpenAI's head of hardware, stated that Jalapeño offers both lower latency and higher throughput, a combination that AI systems typically require a trade-off between. The chip was developed in collaboration with Broadcom and was first announced in June.
Jalapeño was tested using the InferenceX benchmarking platform, comparing its performance against Nvidia's GB200 or GB300 superchips. OpenAI reported that Jalapeño achieved 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower latency.
The chip also demonstrated more tokens per user and throughput per kilowatt than current state-of-the-art inference processors. These results were observed across models such as GPT-OSS 120B, DeepSeek R1, and Kimi K2.5.
The development of Jalapeño specifically targets the efficiency of large language model operations, particularly for multi-step agent tasks. AI agents often call models repeatedly, leading to compounding delays over the course of a longer task.
Jalapeño's design aims to provide higher throughput without increasing response times, which is critical for agents that need to complete many steps in sequence.
OpenAI anticipates that Jalapeño will see initial deployment in small volumes by the end of 2026, with more significant deployment expected in 2027. The company's own AI models assisted in the development process of the chip.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
OpenAI disclosed architectural details and performance metrics for its Jalapeño AI inference processor at the Hot Chips conference. The accelerator, co-developed with Broadcom, aims to outperform Nvidia's GB200 and GB300 in low-latency inference and performance-per-watt metrics. This development indicates OpenAI's continued investment in custom hardware for AI workloads.
OpenAI announced its first custom AI chip, Jalapeño, which analysts state could match or exceed Nvidia's Blackwell-class GPUs in inference efficiency. This development could reduce OpenAI's reliance on Nvidia for inference workloads and impact Nvidia's margins in that growing segment of the AI chip market.
OpenAI announced its in-house Jalapeño ASIC, co-developed with Broadcom, demonstrated up to 1.9 times more throughput per kilowatt and 3.6 times lower latency compared to Nvidia's GB200 and GB300 rack systems in inference benchmarks. This development signifies OpenAI's entry into custom AI hardware, potentially reducing its reliance on external GPU providers for inference workloads.
OpenAI presented benchmark results for its Jalapeño chip at the Hot Chips conference, demonstrating improved tokens per user and throughput per kilowatt compared to current state-of-the-art inference processors. This development indicates progress in specialized hardware for AI inference, potentially leading to more efficient and faster AI model deployment.
OpenAI published initial performance results for its custom Jalapeño inference chip, developed with Broadcom, across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 models. The results indicate higher throughput without increased response times, specifically addressing the compounding delays experienced by AI agents. This development is significant for improving the efficiency of large language model operations, particularly for multi-step agent tasks.
OpenAI announced its new Jalapeño AI chip, developed in partnership with Broadcom, delivers faster responses and higher throughput for AI inference tasks compared to Nvidia's superchips. The chip showed 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower latency in benchmark tests. This development could lead to more responsive AI systems and improved access as demand for AI grows.