Between 2000 and 2004, Intel and AMD raced CPU clock speeds to a thermal wall: the Pentium 4 grew so hot that the industry abandoned the metric entirely and re-architected around multiple slower cores. The gigahertz number was real, marketable and ultimately the wrong axis of competition. OpenAI's new Ultrafast mode — its flagship model served at up to 14 times standard speed — reopens that war on a new substrate, and the question is whether tokens-per-second is the next gigahertz: a metric that matters until it suddenly doesn't.

The Announcement, Stripped of Adjectives

OpenAI this week previewed Ultrafast, a service tier that runs its GPT-5.6 Sol model up to 14 times faster than standard processing, delivering roughly 750 output tokens per second, initially to a select group of API customers. The tier is powered by Cerebras wafer-scale hardware, and the two companies published a joint technical preview of the acceleration stack.

"Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed." — OpenAI, in its official announcement. Source

Three Shifts the Speed Number Hides

First, latency is becoming a priced dimension of intelligence, creating a two-speed market. At 750 tokens per second, the binding constraint flips from model to human: reading speed, review cycles, approval chains. The product that benefits is not the chatbot but the agentic loop — code-test-debug cycles, trading systems, live voice — where each machine-to-machine round trip compounds. Ultrafast is therefore less a faster chat and more the opening price of autonomous-workflow infrastructure.

Second, the Cerebras partnership is a supply-chain maneuver wearing a product announcement's clothes. By proving a frontier model can run at production scale on wafer-scale silicon, OpenAI creates a credible second source for inference and reopens its negotiating position with Nvidia from a posture of monopsony dependence. The speed metric is the marketing; the silicon diversification is the strategy.

Third, the token economy's cost curve is bending. When thinking time approaches zero cost, agent architectures will spend more tokens on internal deliberation than on output, and the industry's unit of account will drift from tokens-generated to tasks-completed. Pricing pages that bill per token will look, within a few cycles, like telephone companies billing per minute of a call that no longer takes minutes.

The Batch-Workload Counter

The skeptic's ledger is straightforward: most enterprise volume is batch, asynchronous and price-sensitive, which is exactly why OpenAI simultaneously cut its cheapest model's price by 80 percent. The revenue center of gravity sits in the discount tier; the latency tier is a halo product for demos and a handful of high-frequency use cases. Speed, on this reading, is a press release with an API attached.

The Prescott Precedent

The historical rhyme is Intel's Pentium 4 Prescott, the 2004 chip that pushed the clock-speed metric to its thermal limit and forced the industry's pivot to multicore and, eventually, to the specialized accelerators that made today's AI possible. The lesson is not that speed races are fake; it is that single-metric races end in physics, and the winners are the firms that re-architect around efficiency before the wall arrives. Ultrafast's 14x will hit its own wall — memory bandwidth, power, cost-per-task — and the companies that build for cost-per-completed-task rather than tokens-per-second will inherit the next cycle.

The Vendor-Claim Counter

A second caution: 14x is a peak figure under favorable conditions, and sustained throughput in production depends on batching, network and memory behavior that previews do not disclose. And while Cerebras reduces Nvidia dependence, it replaces one concentration with another: a single-supplier wafer-scale stack carries its own yield and capacity risks. Diversification between two monopolies is not yet a market.

Positioning for the Latency Market

  • Developers: benchmark agentic loops on cost-per-completed-task, not tokens-per-second; the 14x figure will not survive your queueing delay.
  • Businesses with cycle-time revenue — trading, customer voice, live coding assistance — should pilot the fast tier now, while preview pricing is promotional.
  • Everyone else: stay on the discounted batch tiers; the 80 percent cut is the economically significant announcement of the fortnight.
  • Citizens: expect your assistant to feel instant and your bill to acquire a new premium line; speed tiers migrate from API to consumer within two product cycles.

February 2027: The Speed Tier Becomes Table Stakes

Within six months, expect rival labs to ship 500-plus token-per-second tiers, Nvidia to answer with latency-optimized SKUs, and the fast tier to become the default for paid consumer plans while the price-per-token on speed tiers collapses by half. The gigahertz war took four years to hit its thermal wall; the tokens-per-second war will move faster, because this time every participant has read the history.

Primary source trail: OpenAI's official Ultrafast preview post and Cerebras' joint technical preview.