Lead story
OpenAI's Jalapeño Chip Is the Quiet Revolution Hiding Behind Every ChatGPT Response
For the past few years, running OpenAI's models has meant renting Nvidia's hardware at eye-watering cost. That arrangement just got a lot shakier. Independent benchmarking firm SemiAnalysis has released results from its InferenceX benchmark showing OpenAI's custom silicon — codenamed Jalapeño — beats the current state of the art on two metrics that actually matter: tokens delivered per user, and throughput per kilowatt.
That second number is the one CFOs care about. Inference — the process of running a trained model to generate an answer — is now OpenAI's dominant cost. Training a model is expensive once. Running it billions of times a day is expensive forever. A chip that does more work per watt of electricity isn't just a performance trophy; it's a structural margin advantage that compounds every time someone asks ChatGPT a question.
Why this changes the competitive landscape. Google has its TPUs, Amazon has Trainium, and now OpenAI has Jalapeño. The common thread is that hyperscalers who run AI at massive scale all eventually decide Nvidia's general-purpose GPUs — brilliant as they are — leave efficiency gains on the table. Custom silicon lets you prune exactly for the workload you run most. OpenAI runs inference, all day, every day. Jalapeño is shaped around that specific job.
The benchmark results also put pressure on Anthropic, Mistral, and any other frontier lab that remains fully dependent on external hardware. If Jalapeño gives OpenAI a meaningful cost-per-token advantage, they can cut prices without sacrificing margin — something rivals would struggle to match.
What we still don't know. SemiAnalysis benchmarked Jalapeño, but OpenAI hasn't confirmed when or whether the chip will see broader deployment, or whether it will remain purely internal. There's a non-trivial precedent — Google eventually offered TPU access to cloud customers — but OpenAI's commercial model is different enough that a public cloud play isn't obvious.
The Nvidia angle is also worth watching. Nvidia's share price has ridden the AI wave largely on the assumption that nobody else could build competitive inference silicon at scale. Jalapeño doesn't dethrone Nvidia overnight — training still runs on GPUs, and the installed base is enormous — but it's a credible data point that the moat is narrowing at the edges.
For Australian readers, the cost-of-inference question has direct relevance: every Australian organisation building on OpenAI's API is, in effect, a downstream beneficiary if Jalapeño drives prices down. The federal government's AI in Government Framework and several state-level AI procurement policies have flagged cost and reliability of frontier AI access as key concerns. Cheaper, faster inference makes the policy calculus easier — and the vendor lock-in calculus harder.
Watch for OpenAI to use Jalapeño as quiet leverage in enterprise contract negotiations well before any public announcement. The chip's existence is now confirmed. The question is how aggressively they use it.
