共有:

Inference Hardware

Drop the GPU,
and replies get 14x faster.

OpenAI has previewed a new API tier called "Ultrafast." By moving GPT-5.6 Sol onto Cerebras's dedicated inference chips, response speed reaches up to 14x standard, hitting 750 tokens per second.

AI Navigate Editorial2026.08.157 min read

GPT-5.6 standard Opus 4.8 Fast Claude Fable 5 GPT-5.6 Sol / Ultrafast (Cerebras) up to 750 tok/s · 14x
01
The Hardware

From racks of GPUs
to a single wafer

Ultrafast's speed comes not from a smarter model, but from a different foundation underneath it.

On August 13, 2026, OpenAI announced on its blog a preview of the new "Ultrafast" API tier. Where a standard GPU cluster splits a model across racks of chips, Ultrafast runs on Cerebras's "Wafer-Scale Engine." Cut whole from a single, dinner-plate-sized silicon wafer, the chip carries 44 GB of SRAM on-die and can hold an entire model on one chip — sidestepping the memory-bandwidth bottleneck typical GPU setups run into, according to Cerebras's press release.

The result: GPT-5.6 Sol on Ultrafast reaches up to 750 tokens per second, up to 14x the standard response speed. This isn't OpenAI's first break from GPU-only inference either — it shipped its own Broadcom-co-developed chip, "Jalapeño," back in June 2026.

02
Why It Matters Now

What the numbers say
about "feel"

Completion time on a brutal benchmark, Humanity's Last Exam, translated directly into a felt difference.

14x
max speed vs. standard GPT-5.6
11h 11m
time to finish all 2,500 HLE questions
78h 27m
Claude Fable 5's time on the same test

Ultrafast mode reportedly finished all 2,500 questions on Humanity's Last Exam (HLE) in 11 hours 11 minutes. Claude Fable 5 took 78 hours 27 minutes on the same test — nearly a 7x gap. OpenAI's own comparisons also put it 5x faster than Claude Opus 4.8's Fast mode and 11x faster than Claude Fable 5. The benchmark score itself didn't change — what changed is how fast the same answer arrives, and that's the real story here.

03
Who It Helps

Who benefits, and how

For engineers, the biggest wins come in interactive use cases — coding assistance, customer support — where response latency directly shapes the user experience. OpenAI says customers are already running production tests during the preview across coding, commerce, financial research, and support workflows.

For batch-processing workflows (overnight bulk summarization, non-interactive data processing), total throughput and cost matter more than per-response latency, so the 14x figure delivers little practical benefit there. It's also worth remembering this is a limited, invite-only preview — not something every user can access right now.

Speed doesn't replace
intelligence
it just removes a reason to wait.

04
Next Steps

What to do next

01

Check whether you have access

It's a limited preview — start by checking OpenAI's announcement or your API dashboard for eligibility.

02

Test interactive use cases first

Prioritize places where response latency is the actual bottleneck — coding assistants, support chat — over batch workloads.

03

Check pricing and rate limits separately

Pricing and usage caps may differ from the standard tier behind the speed gains. Confirm the cost structure before production rollout.


05
Risks & Limits

The other side of
the Cerebras dependency

Ultrafast is still a limited preview — general availability and final pricing aren't confirmed yet. This speed also rests on dependency on a single external chip vendor, Cerebras. With multiple AI companies — GPT-5.6 Sol and Devin's SWE-1.7 among them — converging on Cerebras around the same time, the race for speed is, in effect, also a race toward concentrated dependency on one hardware vendor. That's worth keeping in view alongside the headline number.