OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

Summarized from techcrunch.com


At the Hot Chips conference, OpenAI presented detailed benchmark results for its newly developed Jalapeño chip, which was tested using SemiAnalysis’ InferenceX benchmark. According to the benchmarks, Jalapeño demonstrated superior performance compared to current state-of-the-art inference processors, achieving higher tokens per user and greater throughput per kilowatt. OpenAI’s head of hardware, Richard Ho, emphasized that Jalapeño represents a “very, very significant performance advance,” offering both high efficiency in serving numerous customers and low latency in response times.

Jalapeño, first announced in October of the previous year, was developed through a close collaboration between OpenAI and Broadcom, with OpenAI’s own models contributing to the development process. The chip is envisioned as a multigenerational platform, integrating AI products, models, chips, and memory developed in concert. OpenAI’s full-stack approach allowed the company to target specific phases in the inference process that often cause bottlenecks, particularly the prefill and communication phases. Jalapeño is designed to minimize data movement and communication delays by keeping the model state, including the KV cache, local while dynamically activating the appropriate combination of compute, memory, and networking resources for each inference phase. Deployment of Jalapeño is expected to begin in limited volumes by the end of 2026, with more substantial deployment anticipated in 2027.

Source