About TrendForce News

TrendForce News operates independently from our research team, curating key semiconductor and tech updates to support timely, informed decisions.

[News] OpenAI Debuts Jalapeño AI Inference Chip, with Samsung Reportedly Supplying HBM4


2026-08-26 Semiconductors editor

Just ahead of NVIDIA’s latest earnings report on August 26, OpenAI unveiled its first custom inference chip, Jalapeño. In initial benchmarks, it delivered 1.5×–1.9× higher throughput per kilowatt and 1.7×–3.6× lower end-to-end latency than NVIDIA’s GB200 and GB300 rack systems, Tom’s Hardware reports. Notably, beyond the performance claims, Jalapeño could also emerge as a significant new contender for HBM4 supply, the report suggests.

Key Suppliers in Focus

According to Tom’s Hardware, each Jalapeño package combines its compute die with six HBM4 stacks, delivering 216 GiB of memory and 15.4 TB/s of bandwidth. By comparison, NVIDIA’s GB300 features 288GB of HBM3e and carries a 1,400W power rating, the report notes, adding that on a rated-power basis, that gives Jalapeno roughly 50% more memory per watt.

Meanwhile, Economy Tribune and TokenPost report that Samsung Electronics is believed to supply the HBM4 for Jalapeno, with OpenAI handling the architecture and Broadcom as its ASIC partner.

That supply could become a key constraint as HBM remains one of the semiconductor industry’s tightest bottlenecks. Samsung, SK hynix and Micron have already committed capacity through 2027, with the squeeze reportedly prompting NVIDIA to test pared-down Rubin Ultra configurations with as little as 192GB of memory, Tom’s Hardware adds.

If OpenAI scales Jalapeño under its 10GW deployment deal with Broadcom, signed last October, the chip could become a major new claimant on HBM4 supply, adding pressure to a market where NVIDIA already holds substantial capacity, the report suggests.

On the other hand, Tom’s Hardware reports that OpenAI’s first-generation Jalapeño uses TSMC’s 3nm process, meaning the ChatGPT maker will remain in the same race for advanced wafers, HBM and cutting-edge packaging capacity as NVIDIA’s Blackwell and Rubin platforms for the foreseeable future.

OpenAI is already accelerating its Jalapeño roadmap, with Bloomberg reporting that the second-generation chip could reach tape-out within months, while development of a third-generation design is already underway.

Jalapeno’s Benchmark Caveats

While Jalapeño delivers strong performance at relatively low power, two caveats stand out. As noted by Bloomberg, it has yet to be benchmarked against NVIDIA’s Vera Rubin processors, which only recently began shipping. Nor is the chip designed for AI training, where NVIDIA remains dominant. Instead, Jalapeno targets inference—the post-training stage when AI models process prompts, generate responses and perform tasks.

Bloomberg, citing OpenAI chip chief Richard Ho, highlights that AI chips typically involve trade-offs between processing power and response speed, but Jalapeño is designed to deliver both. OpenAI will decide which AI models run on the chip, giving customers greater flexibility to prioritize lower costs or higher performance depending on their needs, the report says.

For its public benchmarks, OpenAI tested Jalapeño on one of its smaller open-source models, along with third-party models from DeepSeek and Moonshot AI. The inclusion of DeepSeek R1 and Kimi K2.5 is particularly notable, suggesting Jalapeno is built for more than OpenAI’s own models and can handle a broader range of large-model workloads, 36Kr notes.

According to Tom’s Hardware, Jalapeño delivered 1.5× to 1.9× higher throughput per kilowatt and 1.7× to 3.6× lower end-to-end latency than VIDIA’s GB200 and GB300 rack systems. The 700W chip reportedly achieved those results while going head-to-head with accelerators rated at 1,200W and 1,400W.

OpenAI normalized the benchmark results against each accelerator’s published package TDP, although it said Jalapeno’s sustained power draw during testing remained at or below 550W, the report notes.

Read more

(Photo credit: OpenAI)

Please note that this article cites information from OpenAITom’s HardwareEconomy TribuneTokenPost, Bloomberg and 36Kr.


Get in touch with us