Inference Hardware
Drop the GPU,
and replies get 14x faster.
OpenAI has previewed a new API tier called "Ultrafast." By moving GPT-5.6 Sol onto Cerebras's dedicated inference chips, response speed reaches up to 14x standard, hitting 750 tokens per second.
From racks of GPUs
to a single wafer
Ultrafast's speed comes not from a smarter model, but from a different foundation underneath it.
On August 13, 2026, OpenAI announced on its blog a preview of the new "Ultrafast" API tier. Where a standard GPU cluster splits a model across racks of chips, Ultrafast runs on Cerebras's "Wafer-Scale Engine." Cut whole from a single, dinner-plate-sized silicon wafer, the chip carries 44 GB of SRAM on-die and can hold an entire model on one chip — sidestepping the memory-bandwidth bottleneck typical GPU setups run into, according to Cerebras's press release.
The result: GPT-5.6 Sol on Ultrafast reaches up to 750 tokens per second, up to 14x the standard response speed. This isn't OpenAI's first break from GPU-only inference either — it shipped its own Broadcom-co-developed chip, "Jalapeño," back in June 2026.
What the numbers say
about "feel"
Completion time on a brutal benchmark, Humanity's Last Exam, translated directly into a felt difference.
Ultrafast mode reportedly finished all 2,500 questions on Humanity's Last Exam (HLE) in 11 hours 11 minutes. Claude Fable 5 took 78 hours 27 minutes on the same test — nearly a 7x gap. OpenAI's own comparisons also put it 5x faster than Claude Opus 4.8's Fast mode and 11x faster than Claude Fable 5. The benchmark score itself didn't change — what changed is how fast the same answer arrives, and that's the real story here.
Who benefits, and how
For engineers, the biggest wins come in interactive use cases — coding assistance, customer support — where response latency directly shapes the user experience. OpenAI says customers are already running production tests during the preview across coding, commerce, financial research, and support workflows.
For batch-processing workflows (overnight bulk summarization, non-interactive data processing), total throughput and cost matter more than per-response latency, so the 14x figure delivers little practical benefit there. It's also worth remembering this is a limited, invite-only preview — not something every user can access right now.
Speed doesn't replace
intelligence —
it just removes a reason to wait.
What to do next
Check whether you have access
It's a limited preview — start by checking OpenAI's announcement or your API dashboard for eligibility.
Test interactive use cases first
Prioritize places where response latency is the actual bottleneck — coding assistants, support chat — over batch workloads.
Check pricing and rate limits separately
Pricing and usage caps may differ from the standard tier behind the speed gains. Confirm the cost structure before production rollout.
The other side of
the Cerebras dependency
Ultrafast is still a limited preview — general availability and final pricing aren't confirmed yet. This speed also rests on dependency on a single external chip vendor, Cerebras. With multiple AI companies — GPT-5.6 Sol and Devin's SWE-1.7 among them — converging on Cerebras around the same time, the race for speed is, in effect, also a race toward concentrated dependency on one hardware vendor. That's worth keeping in view alongside the headline number.