Silicon Desk / OpenAI / Jalapeño / Source-backed briefing / 2026-08-25
← Back to The AGI Times
The AGI Times
Source Notes Desk
Editorial still life of a labeled accelerator wafer and brass calipers on a walnut news desk
OpenAI / Jalapeño / Aug 25

OpenAI Published Jalapeño’s First InferenceX Numbers. You Still Cannot Buy The Chip.

OpenAI published its own InferenceX results for Jalapeño on August 25, 2026: 1.5–1.9× work per watt at peak and 1.7–3.6× lower end-to-end latency on the models it named. The chip is not for sale.

Quick answerOn August 25, 2026, OpenAI published Jalapeño’s first InferenceX numbers. Across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T, it claims 1.5 to 1.9 times more AI work per watt at peak and 1.7 to 3.6 times lower end-to-end latency than the comparison systems it used. The chip is rated 700W, with measured sustained power at or below 550W on the tested loads. You still cannot buy it. OpenAI says it plans to deploy Jalapeño in its own infrastructure by the end of the year and will keep using NVIDIA and other partners.

On August 25, 2026, OpenAI published Jalapeño’s first results. The numbers are OpenAI’s own InferenceX measurements, not an independent lab copy. Jalapeño is OpenAI’s first custom inference chip. The post says it can serve more AI work per unit of power and return responses more quickly, and that the architecture is meant to avoid the usual tradeoff between throughput and latency. None of that is a store listing. Alberta operators cannot buy this silicon.

industrial AI Albertaprivate AI security
1.5–1.9× work/wattPeak throughput vs comparison systems, OpenAI’s InferenceX write-up.
1.7–3.6× lower E2EEnd-to-end latency. Interactive workloads: 2.1–4.1×.
700W ratedSustained power at or below 550W on the loads OpenAI tested.
Not for saleDeploy in OpenAI infra by year-end, per OpenAI. Keep NVIDIA.

What OpenAI published on InferenceX

OpenAI says it tested Jalapeño on InferenceX, a public benchmark from SemiAnalysis, across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. Across those three, it reports 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. For highly interactive workloads, it reports 2.1 to 4.1 times higher performance. It normalized using each accelerator’s published chip power rating. Jalapeño is rated at 700 watts; OpenAI says measured sustained power stayed at or below 550 watts on the workloads tested.

The appendix is OpenAI’s table, not this desk’s lab. For GPT-OSS 120B it lists package TDP as Jalapeño 700W and GB200 1,200W, higher peak mixed TPS/kW of about 1.9× (85,448 vs 44,960 mixed/kW), and lower end-to-end latency of about 1.7× (1.03s vs 1.80s). For Kimi K2.5 1T it lists GB300 at 1,400W, about 1.5× higher peak mixed TPS/kW, and about 3.4× lower end-to-end latency. DeepSeek R1 is also in that appendix, with GB300 at 1,400W. Those figures are copied from OpenAI’s page. They are not an independent confirmation.

“Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems.”OpenAI, Jalapeño’s first results, August 25, 2026

OpenAI describes the architecture around real language-model phases: prefill is compute-intensive; decode is constrained more by memory bandwidth; KV cache is meant to stay local. It says AI helped move from initial design to tapeout in nine months. Using Codex with GPT-Astra, it says the team brought three open-weight models that were not in the original production plan to high performance within two months. For selected GPT-OSS attention and mixture-of-experts blocks, it says AI-generated implementations ran 1.5 to 1.8 times faster than existing human-expert kernels. OpenAI says those figures apply to the selected blocks, not the full model.

You still cannot buy the chip

OpenAI plans to begin deploying Jalapeño inside its own compute infrastructure by the end of the year. It calls Jalapeño the first generation of a multigenerational roadmap, with Gen 2 deep in development. It also says it will continue to widely deploy accelerators from NVIDIA and other partners for training and inference. That is a first-party silicon story. It is not a channel SKU for an Edmonton shop.

The June 24 unveil post, OpenAI and Broadcom unveil LLM-optimized inference chip, names Broadcom and Celestica as partners on implementation, boards, racks, and production systems. This desk did not fetch a TechCrunch recap, so it is not adding TCO-versus-Rubin claims, Hot Chips slides, or volume figures that are not on OpenAI’s pages.

What Alberta operators should actually watch

Industrial AI Alberta teams cannot buy Jalapeño. If OpenAI later serves API traffic from this stack, latency and price might move. That is still a closed full stack: models, serving software, chips, memory, networking. It does not replace a private on-device path. Client files, tenders, and plant data still belong on a private AI security route, not on a hope that OpenAI’s watt-per-token chart will someday land in your rack.

Keep a second path. If the job has to stay on the desk, it still belongs on a Mac mini or other private stack. If the job is API work, watch OpenAI’s own latency and price after year-end deployment — and do not treat a first-party InferenceX appendix as a purchase order. AI consulting here means mapping which workflows can wait on OpenAI’s infra and which cannot leave the plant.

The guardrail

OpenAI’s numbers are OpenAI’s numbers. InferenceX is named; the comparison systems and TDP lines are in the appendix. That is not an independent audit, and it is not a chip you can order. For private AI security, a faster closed stack is still someone else’s building. Keep NVIDIA in the plan because OpenAI says it will. Keep on-device work on-device because Jalapeño will not show up in receiving.

Opcelerate recommendationDo not budget to buy Jalapeño. Read OpenAI’s InferenceX appendix as a first-party claim. Watch API latency and price only if OpenAI deploys this in its own infra. Keep a second, private path for work that cannot sit in a closed full stack. Opcelerate can map which jobs wait on that API and which stay on-device.