OpenAI officially unveiled Jalapeño, its first AI accelerator chip, on August 25. The chip is capable of up to 13.4 petaflops of 4-bit compute and features 232 gigabytes of advanced memory, with a bandwidth of 15.4 terabytes per second.
Benchmarks provided by OpenAI indicate that Jalapeño can reduce end-to-end latency by up to 3.6 times compared to Nvidia’s GB300, while also consuming less power. These performance claims are for its use in OpenAI's inference fleet.
A notable aspect of Jalapeño's development is the use of OpenAI's own large language models (LLMs) in its design. This approach contributed to a rapid development cycle, moving from initial architectural concept to first silicon in under 20 months.
The time from the first register-transfer level (RTL) code to tape-out, when the design is sent for manufacturing, was nine months. Richard Ho, OpenAI's vice president of hardware, stated that LLMs provided "superpowers" to engineers, allowing for faster exploration of design paths.
The team responsible for designing Jalapeño averaged fewer than 100 people throughout the project. This team handled the end-to-end system design, including the inference accelerator, memory hierarchy, and networking.
OpenAI partnered with Broadcom, which was responsible for the physical design aspects of the chip, starting from the gates. This division of labor contributed to the project's timeline.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
OpenAI is deploying its new Jalapeño ASICs with AMD EPYC Turin CPUs, each featuring 1.5TB of memory. OpenAI's VP of Hardware, Richard Ho, stated that the decision to use Turin was pragmatic due to Nvidia's Vera CPU being less mature, prioritizing de-risking and faster deployment.
OpenAI designed its Jalapeño ASIC from register-transfer level (RTL) to tapeout in nine months using AI-assisted methods. This accelerated timeline, which is significantly shorter than the industry standard, demonstrates a new baseline for hardware development with AI.
OpenAI's Head of Hardware, Richard Ho, confirmed that the company's custom Jalapeño AI inference ASIC is currently intended for internal use to meet OpenAI's compute demands. However, Ho indicated that a broader rollout of the chip to other users is a possibility in the future. This clarifies the initial purpose of the ASIC following its reveal and benchmarks against Nvidia's accelerators.
OpenAI's VP of Hardware, Richard Ho, discussed the company's Jalapeño inference ASIC, emphasizing efficiency as the primary driver for its development. The chip aims to address compute limitations and power constraints in data centers, focusing on low-latency inference for user-facing AI applications.
OpenAI has introduced Jalapeño, its first AI accelerator chip, which offers up to 13.4 petaflops of 4-bit compute and 232 GB of memory. The chip was designed in under 20 months, with OpenAI's own large language models (LLMs) assisting the design process, enabling a faster development timeline.