OpenAI's Jalapeño Chip Shows Significant AI Inference Gains
On August 25, 2026, OpenAI unveiled benchmark results for its Jalapeño chip, reporting 1.5 to 1.9 times more AI work per watt. This custom inference chip improves peak throughput and reduces latency across models like GPT-OSS 120B and others, promising faster AI responses and broader availability.

On August 25, 2026, OpenAI introduced significant advancements with the release of benchmark results for its custom inference chip, Jalapeño. This new chip has demonstrated impressive performance improvements across various AI models, including GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. The benchmarks indicate that Jalapeño can perform 1.5 to 1.9 times more AI work per watt at peak throughput, while also achieving lower latency, ranging from 1.7 to 3.6 times less than existing systems. This innovation is expected to enhance the efficiency of AI applications, allowing for faster responses and more responsive AI agents as demand continues to grow.
OpenAI aims to ensure that advancements in artificial general intelligence benefit all of humanity, and Jalapeño is a step toward making advanced AI more accessible and affordable. The development of Jalapeño is rooted in OpenAI's historical engagement with AI models. The performance of earlier model generations played a crucial role in the design and optimization of Jalapeño. This chip benefits from a full-stack approach, where OpenAI integrates the design of models, chips, software, and systems to create a cohesive and efficient architecture.
As a result, Jalapeño provides both high throughput and low latency without the trade-offs typically seen in existing hardware systems. This comprehensive architecture allows Jalapeño to excel across various AI workloads, making it a versatile solution for future AI applications. The implementation of Jalapeño involved rigorous testing and benchmarking through InferenceX, a public standard that evaluates the process of serving AI requests. Jalapeño was compared against leading AI systems in terms of power efficiency, throughput, and latency.
The results consistently showcased Jalapeño's superior performance, particularly in high-interaction scenarios, where it demonstrated 2.1 to 4.1 times higher performance than competing systems. By measuring performance per unit of power, OpenAI provides a more relevant standard for assessing AI capabilities, highlighting Jalapeño's competitive edge in the market. The broader implications of Jalapeño's performance extend into various sectors, including technology and industry.
The chip's increased efficiency and responsiveness can lead to advancements in applications such as natural language processing, real-time data analysis, and interactive AI systems. As Jalapeño is deployed in the coming months, it is expected to contribute to the development of more capable AI products, fostering innovation across multiple domains. This shift towards more efficient AI systems can help meet the growing demands of users and businesses alike.
Looking ahead, OpenAI plans to ramp up the production and deployment of Jalapeño, aiming to deliver even faster and more capable AI products. The success of this chip marks the beginning of a multigenerational platform that will evolve with advancements in AI technology. As OpenAI continues to optimize and enhance Jalapeño, its significance in the global landscape of artificial intelligence will likely grow, paving the way for transformative applications that benefit a wide range of industries and communities.
Enjoyed this story?
Show the newsroom a little love — one tap per reader.
Liam covers climate solutions and the people helping the planet heal.
Be part of the good
Stories like this start with people who care. Share it, or submit your own uplifting story to inspire millions today.


