August 25 — OpenAI formally released its first AI accelerator chip, Jalapeño. Jalapeño can deliver up to 13.4 petaflops of 4-bit compute and access 232 GB of state-of-the-art memory at interconnect speeds of up to 15.4 TB/s. Benchmarks cited by OpenAI indicate that, compared with the Nvidia GB300 chips the company currently uses, Jalapeño can reduce end-to-end latency (the time from prompt to last token) by up to 3.6× while also consuming less power.
Whether those figures will translate into real gains once Jalapeño is widely deployed in OpenAI’s inference clusters remains to be seen, but performance is only half the story. The other half is how the chip was designed — and, as you might expect, OpenAI’s large language models (LLMs) accelerated the process. Jalapeño went from initial architectural concept to first silicon in less than 20 months. From the first RTL (register-transfer level code that defines the chip’s logic) to tape-out (when the final design is sent to manufacturing), it took only nine months.
Fast as that timeline is, experts believe it will soon look slow as LLMs improve and become more deeply integrated with chip-design tools. Unsurprisingly, OpenAI is bullish on the opportunity. “These models give our engineers superpowers,” said Richard Ho, OpenAI’s VP of Hardware. “Our engineers are still in the driver’s seat. They’re still the final decision-makers. But they can work faster and explore more paths.”