
They found that the processor is wicked fast compared with «leading chips,» which means Nvidia’s offerings, specifically the Grace Blackwell 200 which they show getting trounced in some tests.
The benchmark found that Jalapeño delivers 1.9x more tokens per second when measured per watt on peak efficiency, and is 17.8x faster on higher token density, which translates to raw performance in handling requests.
They also found that end-to-end latency was between 1.7x and 3.4x lower depending on the model, meaning the user will spend less time waiting for responses.
These two measures combined are key, as most chips have to make tradeoffs between latency and throughput, and few can be good at both, OpenAI hardware vice president Richard Ho tells The Verge.
Jalapeño was developed in record time — 9 to 16 months — assisted by AI, and OpenAI says their upcoming Astra model is already busy working on the second generation, which they say is «in deep development.»
OpenAI will deploy and «operate Jalapeño at scale» within their compute infrastructure «by the end of the year.»
Read more: OpenAI’s report, X post. The Verge, Bloomberg (paywalled) and The Register. Discussion on Hacker News and r/Singularity.