OpenAI has officially unveiled its inaugural custom inference processor, named JalapeƱo, which was developed in close collaboration with Broadcom. Initial published benchmarks presented at the Hot Chips conference indicate that this 700W ASIC surpasses Nvidia’s flagship GB200 and GB300 rack systems.
According to the published data, the custom silicon achieves up to 1.9 times higher throughput per kilowatt and up to 3.6 times lower end-to-end latency. These performance disclosures arrive as the technology sector closely watches how hardware velocity shifts the competitive boundaries of optics articles and advanced computing infrastructure.
Architecture and Benchmark Performance
The evaluation benchmarks utilized three prominent open modelsāGPT-OSS 120B, DeepSeek R1 670B, and Moonshot AI’s Kimi K2.5āevaluated via SemiAnalysis’s InferenceX suite. Unlike hardware geared toward generalized training, JalapeƱoās architecture is engineered exclusively for AI inference.
Maximizing Memory Bandwidth
The core design principle relies on maximizing aggregate High Bandwidth Memory utilization to eliminate bottlenecks. Each JalapeƱo package pairs a primary compute die with six HBM4 stacks, delivering 216 GiB of high-speed memory running at 15.4 TB/s.
For enthusiasts tracking hardware advancements, comparing these massive computational leaps brings to mind the precision engineering found in fine microscopes or long-range telescopes. Every microscopic component must align perfectly to achieve optimal signal clarity and processing speed.
Future Deployment and Industry Impact
OpenAI plans to begin deploying these groundbreaking inference processors within its proprietary data centers later this year. Despite this bold move into custom silicon, the organization remains tightly linked to Nvidia hardware through recent infrastructure financing agreements.
Concurrently, development on a second-generation JalapeƱo processor is already advancing rapidly toward tapeout on an advanced TSMC 3nm-class process. As the industry scales toward new horizons, staying informed on these hardware developments ensures professionals understand the future of next-generation infrastructure.
Here is the source article for this story: OpenAIās 700W JalapeƱo ASIC outpaces 1,400W Nvidia flagship GPU ā claims up to 1.9x throughput per kilowatt and 3.6x lower latency, co-developed with Broadcom