OpenAI's Jalapeño chip reports major gains in AI inference
OpenAI has published its first measured results for Jalapeño, a custom inference chip designed for serving large language models. In tests using the public InferenceX benchmark, OpenAI says Jalapeño delivered 1.5–1.9× more AI work per watt and 1.7–3.6× lower end-to-end latency than the commercial comparison systems.
The result matters at the serving layer. OpenAI is not presenting Jalapeño as a general-purpose replacement for training accelerators, and the chip is not available for customers to buy. The immediate significance is narrower and more practical: if the measurements hold at production scale, custom inference hardware could improve response speed and reduce the power required for high-volume agent workloads.
What OpenAI measured on Jalapeño
OpenAI compared Jalapeño with leading commercial systems across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. The company reports that the chip stayed on the performance-per-watt and latency Pareto frontier across the tested operating range, rather than optimizing one metric at the expense of the other.
The published figures are benchmark results, not a customer SLA. Jalapeño is rated at 700 watts, while measured sustained power was at or below 550 watts in the tested workloads. On Kimi K2.5, OpenAI reports about 1.5× higher peak performance per watt and 3.4× lower end-to-end latency than the comparison system. Data Center Dynamics independently reported the same broad performance claims and the 700-watt rating.
Why inference hardware is the strategic target
Inference has different bottlenecks from model training. Prompt processing is compute-intensive; token generation is more constrained by memory bandwidth and data movement. OpenAI says Jalapeño keeps model state, including the KV cache, local within a connected system and coordinates compute, memory, and networking around both phases.
That design is aimed at interactive software. In an agentic workload, a task may require many model calls in sequence, so small latency improvements can compound across the entire run. More work per kilowatt also matters because serving demand grows with every product, API workflow, and coding agent that uses the model.
Axios described the announcement as part of a broader move by AI companies to design their own silicon to reduce cost pressure and dependence on Nvidia. OpenAI says it will continue deploying Nvidia and other partner accelerators for both training and inference, so Jalapeño is an additional layer in the stack rather than an immediate replacement.
The limits of the result
The benchmark is based on OpenAI’s measurements using InferenceX, a public benchmark from SemiAnalysis. The comparisons use published chip power ratings and selected operating points; they do not establish total system cost, availability, cooling requirements, software maturity, or performance across every model and workload.
OpenAI also says that supporting each model family still requires new kernels and model-specific optimization. It reports that Codex with GPT-Astra brought three open-weight models to high performance within two months, and that selected GPT-OSS attention and mixture-of-experts blocks ran 1.5–1.8× faster than existing human-written implementations. Those are selected blocks, not full-model results.
Deployment is planned for late 2026
OpenAI plans to begin deploying Jalapeño inside its own compute infrastructure by the end of 2026. The company describes the chip as the first generation of a multi-generation roadmap, with a second generation already in development. Until production qualification, software maturation, and scale validation are complete, the published numbers should be read as an infrastructure signal rather than a confirmed change to ChatGPT, Codex, or API pricing.
For developers, the next useful checkpoint is operational: whether OpenAI reports production deployment, sustained service metrics, or concrete customer-facing changes. Those details will determine whether Jalapeño changes the economics users actually see, rather than only the benchmark economics of the underlying serving stack.
Sources
Try the related loot
Rent Out Your Idle GPU on Vast.ai—The 75% Revenue Share Comes With Real Work
