NewsStocksOpenAI says its Jalapeño chip beats Nvidia's GB300 on inference

OpenAI says its Jalapeño chip beats Nvidia's GB300 on inference

Author: Cryptopolitan·

Key Takeaways

  • OpenAI said Jalapeño was tested on SemiAnalysis’s InferenceX benchmark using GPT-OSS 120B, DeepSeek’s R1 670B, and Moonshot AI’s Kimi K2.5 1T.
  • The company reported that Jalapeño delivered 1.5 to 1.9 times more work per unit of electricity and 1.7 to 3.6 times faster responses than the comparison systems.
  • OpenAI said the chip’s advantage rose to 2.1 to 4.1 times on tasks that involved more back-and-forth interaction.
  • OpenAI said it does not plan to sell or rent Jalapeño, and that its main constraint is data center power rather than money or floor space.
  • The chip is expected to begin limited deployment by the end of 2026 and ramp through 2027, while OpenAI continues to rely on Nvidia and Cerebras alongside its own silicon.
OpenAI says its Jalapeño chip beats Nvidia's GB300 on inference

OpenAI has published its first benchmark results for Jalapeño, its in-house inference chip built with Broadcom.

The company says the chip does more AI work per watt and delivers faster answers than Nvidia’s GB200 and GB300 rack systems, which were the comparison systems used in the tests.

With Jalapeño, OpenAI joins other large AI operators — Google with its Tensor Processing Units, Amazon with its Trainium chips, Microsoft with its Maia accelerators, and Meta with its MTIA — in designing its own AI silicon as inference costs mount across the industry.

How OpenAI says Jalapeño performed

OpenAI said Jalapeño was evaluated using InferenceX, a public benchmark from SemiAnalysis that measures the full process of serving an AI request. The chip was tested on OpenAI’s GPT-OSS 120B, DeepSeek’s R1 670B, and Moonshot AI’s Kimi K2.5 1T.

According to OpenAI, Jalapeño completed 1.5 to 1.9 times more work per unit of electricity than the other systems in the benchmark. It also returned responses 1.7 to 3.6 times faster.

For tasks that require more back-and-forth interaction, OpenAI said the advantage widened to between 2.1 and 4.1 times.

OpenAI hardware vice president Richard Ho told reporters that the chip can handle more work at once while also responding more quickly, something he said most chips cannot do simultaneously.

The comparison targets were Nvidia’s GB200 and GB300 superchips, which were the strongest results InferenceX had recorded at the time. On the largest model tested, Kimi K2.5, OpenAI said Jalapeño reached about 1.5 times peak performance per watt and 3.4 times lower latency.

OpenAI also said it has no rental market, no instance type, and no plan to sell Jalapeño. That differs from cloud providers such as Amazon and Google, which rent their own custom chips to outside customers. When asked whether the company would offer the chip to others, Ho said OpenAI was “struggling to have enough” compute for its own needs.

Ho said OpenAI has the budget and the floor space it needs, but is constrained by data center power rather than money or physical capacity, making tokens per megawatt the key metric. Power availability has become a gating factor for data center construction industry-wide, with grid capacity increasingly shaping how quickly AI infrastructure can be built. Because inference runs continuously across ChatGPT and the API, any reduction in watts per token compounds at OpenAI’s scale.

The chip is rated at 700 watts, but OpenAI said it was held at or below 550 watts during the workloads it tested.

The hardware also fits within OpenAI’s October 2025 deal with Broadcom to deploy 10 gigawatts of OpenAI-designed accelerators through 2029. Cryptopolitan previously reported when the partnership was first outlined.

OpenAI’s deployment plans

Speaking at Hot Chips, an annual semiconductor engineering conference, Ho said OpenAI will deploy Jalapeño in racks of 128 chips, with a full pod using 2,048 ASICs. He said a 128-chip deployment delivers 1.7 exaflops of 4-bit compute and 27.5 terabytes of HBM4, with each package providing 15.4 terabytes per second of memory bandwidth.

OpenAI said it used its own models to accelerate development, moving from design to tape-out in nine months, a pace well inside the multi-year cycles typical of chip development. Ho also said a second-generation version is “deep into development” and that work on a third version is already underway.

Even so, OpenAI is not replacing its GPU fleet. Ho described Jalapeño as one part of a broader compute strategy that still depends on “very, very good partners” at Nvidia and Cerebras, the wafer-scale chipmaker.

The chip is expected to be deployed in small volumes by the end of 2026 before ramping through 2027.

OpenAI said all of the figures came from the company itself. SemiAnalysis verified the InferenceX runs in person, but did not run the full suite or see results from AgentX, the companion SemiAnalysis benchmark for agentic, multi-step workloads.