OpenAI presents Jalapeño chip and integrated compute strategy

OpenAI published measured results for Jalapeño, its first custom inference chip, reporting higher peak throughput per kilowatt and lower token latency on the InferenceX benchmark using GPT‑OSS 120B. The company said Jalapeño also performed strongly on DeepSeek R1 and Kimi K2, and that future chip generations are already underway.

Performance claims and benchmark details

On InferenceX with GPT‑OSS 120B, OpenAI reported that Jalapeño delivered greater mixed-token throughput per kW and reduced token latency compared with the commercial systems included in the comparison. The company published benchmark figures such as 535.28 tokens/s per user for GPT‑OSS and noted strong per-kilowatt results on DeepSeek R1 and Kimi K2.

OpenAI described the chip as part of an integrated approach: co-designing model, serving software, memory, network and silicon to improve throughput, latency, energy efficiency and cost together, rather than optimizing each layer independently.

Portfolio approach and data-center work

The company said Microsoft’s compute and NVIDIA’s chips have been foundational to its growth, and listed other partners in its compute portfolio: AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy and SoftBank. OpenAI characterized this mix as a way to meet different workload requirements — training, high-volume inference and always-on agents — by choosing premium systems where capability matters and optimizing for efficiency where scale and cost are primary concerns.

OpenAI also highlighted Project Camellia, a Georgia data-center effort it described as tailored to customer workloads. The company said the project includes closed-loop water conservation, commitments to cover infrastructure and energy costs, support for local jobs and businesses, and an annual independent public audit of those commitments.

Efficiency, economics and agent performance

OpenAI cited the Artificial Analysis Coding Agent Index, reporting that GPT‑5.6 Sol with max reasoning reached a new high while using 54% fewer output tokens than another leading model. The company framed such gains as reducing retries and total cost for successful work, and invoked Jevons paradox to note that greater efficiency can expand overall consumption and use cases.

OpenAI published these results and strategic details on August 25, 2026, and stated that subsequent Jalapeño generations are in development.


Original source: OpenAI News

Leave a Comment