OpenAI unveils Jalapeño, its first chip — 700W beat Nvidia parts rated at 1,200W
OpenAI unveiled its first custom AI chip, Jalapeño, on August 25, 2026. It handles inference only — it runs finished models rather than training them. The company reports 1.5 to 1.9 times more AI work per watt at peak throughput across all three models tested, 1.7 to 3.6 times lower end-to-end latency than the best commercially available systems, and 2.1 to 4.1 times higher performance on interactive workloads. On a public benchmark a 700-watt Jalapeño part outperformed Nvidia accelerators rated at 1,200 and 1,400 watts. Analysts noted the comparison is incomplete because Jalapeño uses newer HBM4 memory. Deployment in OpenAI's own infrastructure is targeted for the end of 2026, and the company says it will continue buying Nvidia silicon
The three lines
- What it is — A general-purpose LLM inference accelerator. It runs models; it does not train them
- Numbers — 1.5–1.9x work per watt, 2.1–4.1x on interactive workloads. 700W beat 1,200W and 1,400W parts
- Caveat — Jalapeño uses HBM4. Analysts say the fair comparison is Nvidia's Vera Rubin, which also uses HBM4
Key questions
- What is OpenAI's Jalapeño chip?
- **OpenAI's first internally designed AI chip, and it does inference only.** AI silicon serves two distinct jobs: **training**, which builds a model, and **inference**, which runs a finished model to produce answers. Jalapeño does the second. It was presented not as a chip tuned to one model but as a **general-purpose LLM inference accelerator**. The published figures: across **all three models tested**, **1.5 to 1.9 times more work per watt** at peak throughput, **1.7 to 3.6 times lower** end-to-end latency, and **2.1 to 4.1 times higher** performance on interactive workloads — the real-time back-and-forth a person experiences as chat. The most striking comparison is on power: on a public benchmark a **700-watt Jalapeño part** outperformed Nvidia accelerators rated at **1,200 and 1,400 watts**.
- Does this mean Nvidia has been beaten?
- **Not on this benchmark alone, because of memory.** Jalapeño uses **HBM4**, while the Nvidia products it was measured against use an earlier memory generation. Analysts called the comparison somewhat incomplete and unfair on that basis, and said the appropriate opponent is **Nvidia's Vera Rubin, which also uses HBM4**. There is a second limit: Jalapeño **only does inference**. Training frontier models remains the work of large general-purpose GPU clusters, and that market is a substantial part of Nvidia's revenue. Consistent with that, OpenAI said it **will keep buying Nvidia chips**. This page covered the memory generations in "HBM4 versus HBM3E: the structural change behind double the bandwidth."
- Why is this a significant story?
- **Because an AI company moved from buying silicon to designing it.** The axis of AI competition has been which lab builds the smartest model. That axis is widening to **chips, memory, data centres and electricity**, and Jalapeño is a marker of the shift. The economics are direct: inference means **running one finished model hundreds of millions of times**, so electricity is a recurring unit cost rather than a one-off. Work per watt of 1.5 to 1.9 times means the same electricity does 1.5 to 1.9 times the work, and at scale that difference compounds. The timing sharpened the point — one day later Nvidia reported **$96.2 billion in quarterly revenue with $108 billion guidance**, a record quarter, and the stock fell after hours. Covered separately in "Nvidia Q2 revenue $96.2 billion."
OpenAI unveiled its first in-house AI chip, Jalapeño, on August 25, 2026 local time.
The chip does not train models. It only runs them — inference.
And within that narrow job, one number stands out: a 700-watt part beat a 1,200-watt one.
1. What OpenAI published
| Measure | Reported | Compared against |
|---|---|---|
| Work per watt (peak throughput) | 1.5–1.9× | best commercially available systems |
| End-to-end latency | 1.7–3.6× lower | best commercially available systems |
| Interactive workload performance | 2.1–4.1× | best commercially available systems |
| Public benchmark power | 700W | Nvidia 1,200W and 1,400W |
The company says these results held across all three models tested.
2. Training and inference — why build only for the second
| Training | Inference | |
|---|---|---|
| Job | builds the model | runs the model |
| Frequency | a few times per model | every user request |
| Needs | huge clusters, flexibility | throughput, latency, efficiency per watt |
| Cost type | closer to development spend | closer to unit cost |
The distinguishing feature of inference is repetition. A model is built once and then runs hundreds of millions of times.
That makes every watt a line item. Work per watt of 1.5 to 1.9 times means the same electricity does 1.5 to 1.9 times the work, and at data-centre scale that gap becomes a cost gap.
It also explains why an inference chip is a more tractable engineering target than a training GPU. A narrow job can be designed for narrowly. This page covered the split in "Training versus inference: why the same GPU sells twice," and the chip-design side in today's "What an AI inference chip is."
3. The comparison carries a caveat
Analysts flagged memory generation immediately.
| Item | Jalapeño | Nvidia parts compared |
|---|---|---|
| Memory | HBM4 | earlier generation |
| Fair opponent | — | Vera Rubin (HBM4) |
HBM is the binding constraint in AI silicon. Compute performance is irrelevant if data cannot reach it fast enough, and HBM4 roughly doubles bandwidth over the prior generation. Two chips with the same design will diverge sharply if their memory generations differ.
Analysts therefore described the published comparison as somewhat incomplete and unfair, and named Vera Rubin — also HBM4 — as the proper benchmark opponent.
One more limit: these are OpenAI's own measurements. No independent reproduction has been published.
4. Buying and building at the same time
The odd part of the announcement is its conclusion. OpenAI unveiled a chip of its own and said it will keep buying Nvidia's.
That is not a contradiction once the work is split:
- Training stays purchased — frontier model training remains a large-GPU-cluster problem.
- Inference gets built in-house — the repeated, high-volume cost centre moves onto custom silicon.
- Timing — deployment in OpenAI's own infrastructure is targeted for end of 2026.
CNBC framed the result as a threat to Nvidia's margins rather than to its market. The point is not that Nvidia loses customers; it is that the highest-volume, most repetitive slice moves in-house, and the price Nvidia commands on that slice compresses.
5. Nvidia's answer, one day later
On August 26, Nvidia reported fiscal 2027 second-quarter results.
| Item | Figure |
|---|---|
| Q2 revenue | $96.2bn (+106% YoY) |
| Data Center | $89.02bn (+117% YoY) |
| Q3 guidance | $108.0bn ±2% |
| After-hours stock | about -1.3% |
A record quarter, guidance above consensus, and a lower stock. That these two stories sit one day apart is itself a description of what the AI silicon market is currently worried about.
6. What could not be confirmed
- Announcement date — sources give August 25 or 26; the former is used here.
- Test conditions — the three models, the benchmark name and its conditions are unconfirmed.
- Comparison parts — only the 1,200W and 1,400W ratings are confirmed, not model names.
- Manufacturing — foundry and process node unknown.
- Independent verification — no third-party reproduction of the performance figures.
- Next checkpoint — actual deployment at the end of 2026. Whether the benchmark numbers reproduce inside OpenAI's own data centres is the real test.
Sources
- CNBC — OpenAI's Jalapeño AI chip brings new 'threat' to Nvidia margins as custom silicon gains ground
- The Decoder — OpenAI's first custom chip "Jalapeño" reportedly beats Nvidia's Blackwell and Rubin in inference benchmarks
- SemiAnalysis — OpenAI Jalapeño: Better Than Nvidia Blackwell
- 24/7 Wall St. — OpenAI's Custom Chip Embarrasses Nvidia, While Company Vows to Keep Buying From It
- ForkLog — OpenAI says Jalapeño outperforms Nvidia Blackwell in tests
- Stocktwits — OpenAI's Jalapeno Chip Is Outperforming Nvidia, AMD And Google Chips, SemiAnalysis Says