← All stories
● Covered by 5 sources · 6 reportsMedium impact6 neutral

OpenAI's Jalapeño AI Chip Shows Faster Responses and Higher Throughput in Benchmarks

🔄 Updated 7d ago — new reporting from Tom's Hardware
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Jalapeño is an Application-Specific Integrated Circuit (ASIC) for AI inference.
  • Developed with Broadcom, it was first introduced in June.
  • Outperformed Nvidia's GB200/GB300 in InferenceX benchmarks.
  • Showed 1.5-1.9x more AI work per watt and 1.7-3.6x lower latency.
  • Deployment is expected in small volumes by late 2026, with more in 2027.
  • Jalapeño is a 700W part, compared to Nvidia's 1,200W and 1,400W accelerators.
  • OpenAI plans to deploy the chip in its own data centers later this year.
  • Tests covered GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI's Kimi K2.5 models.
  • Jalapeño showed 8.6 to 104.3 times more throughput per kilowatt at low-latency operating points.
  • Jalapeño's measured sustained power stayed at or below 550W in testing.
  • OpenAI unveiled the Jalapeño semiconductor on Tuesday.
  • Jalapeño reached tape-out in nine months.
  • Jalapeño uses a NUMA-style spatial architecture.
  • A 2,048-processor system scales to 27 exaFLOPS and 32 PB/s of aggregate memory bandwidth.
  • Jalapeño has 216 GB of HBM4 memory with up to 15.4 TB/s bandwidth.
  • Jalapeño delivers up to 3.4 MXFP8 PFLOPS and 13.4 MXFP4 PFLOPS.
  • Jalapeño operates at 1.70 GHz, with plans to increase to 1.80 GHz.

OpenAI Releases Jalapeño Benchmarks

OpenAI has published initial performance results for its custom Jalapeño AI inference chip. The benchmarks, presented at the Hot Chips conference, indicate that Jalapeño delivers faster responses and higher throughput for AI inference tasks.

Richard Ho, OpenAI's head of hardware, stated that Jalapeño offers both lower latency and higher throughput, a combination that AI systems typically require a trade-off between. The chip was developed in collaboration with Broadcom and was first announced in June.

Performance Against Competitors

Jalapeño was tested using the InferenceX benchmarking platform, comparing its performance against Nvidia's GB200 or GB300 superchips. OpenAI reported that Jalapeño achieved 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower latency.

The chip also demonstrated more tokens per user and throughput per kilowatt than current state-of-the-art inference processors. These results were observed across models such as GPT-OSS 120B, DeepSeek R1, and Kimi K2.5.

Addressing AI Agent Needs

The development of Jalapeño specifically targets the efficiency of large language model operations, particularly for multi-step agent tasks. AI agents often call models repeatedly, leading to compounding delays over the course of a longer task.

Jalapeño's design aims to provide higher throughput without increasing response times, which is critical for agents that need to complete many steps in sequence.

Deployment Timeline

OpenAI anticipates that Jalapeño will see initial deployment in small volumes by the end of 2026, with more significant deployment expected in 2027. The company's own AI models assisted in the development process of the chip.

Updates

🕒 2026-08-27 · new reporting from Tom's Hardware
  • Jalapeño reached tape-out in nine months.
  • Jalapeño uses a NUMA-style spatial architecture.
  • A 2,048-processor system scales to 27 exaFLOPS and 32 PB/s of aggregate memory bandwidth.
  • Jalapeño has 216 GB of HBM4 memory with up to 15.4 TB/s bandwidth.
  • Jalapeño delivers up to 3.4 MXFP8 PFLOPS and 13.4 MXFP4 PFLOPS.
  • Jalapeño operates at 1.70 GHz, with plans to increase to 1.80 GHz.
🕒 2026-08-26 · new reporting from CNBC Technology
  • OpenAI unveiled the Jalapeño semiconductor on Tuesday.
🕒 2026-08-25 · new reporting from Tom's Hardware
  • Jalapeño is a 700W part, compared to Nvidia's 1,200W and 1,400W accelerators.
  • OpenAI plans to deploy the chip in its own data centers later this year.
  • Tests covered GPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI's Kimi K2.5 models.
  • Jalapeño showed 8.6 to 104.3 times more throughput per kilowatt at low-latency operating points.
  • Jalapeño's measured sustained power stayed at or below 550W in testing.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~19 min · 16 stories · Sep 04

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

How outlets covered it

OpenAI disclosed architectural details and performance metrics for its Jalapeño AI inference processor at the Hot Chips conference. The accelerator, co-developed with Broadcom, aims to outperform Nvidia's GB200 and GB300 in low-latency inference and performance-per-watt metrics. This development indicates OpenAI's continued investment in custom hardware for AI workloads.

OpenAI announced its first custom AI chip, Jalapeño, which analysts state could match or exceed Nvidia's Blackwell-class GPUs in inference efficiency. This development could reduce OpenAI's reliance on Nvidia for inference workloads and impact Nvidia's margins in that growing segment of the AI chip market.

OpenAI announced its in-house Jalapeño ASIC, co-developed with Broadcom, demonstrated up to 1.9 times more throughput per kilowatt and 3.6 times lower latency compared to Nvidia's GB200 and GB300 rack systems in inference benchmarks. This development signifies OpenAI's entry into custom AI hardware, potentially reducing its reliance on external GPU providers for inference workloads.

OpenAI presented benchmark results for its Jalapeño chip at the Hot Chips conference, demonstrating improved tokens per user and throughput per kilowatt compared to current state-of-the-art inference processors. This development indicates progress in specialized hardware for AI inference, potentially leading to more efficient and faster AI model deployment.

OpenAI published initial performance results for its custom Jalapeño inference chip, developed with Broadcom, across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 models. The results indicate higher throughput without increased response times, specifically addressing the compounding delays experienced by AI agents. This development is significant for improving the efficiency of large language model operations, particularly for multi-step agent tasks.

OpenAI announced its new Jalapeño AI chip, developed in partnership with Broadcom, delivers faster responses and higher throughput for AI inference tasks compared to Nvidia's superchips. The chip showed 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower latency in benchmark tests. This development could lead to more responsive AI systems and improved access as demand for AI grows.