NewsMacroGoogle Ships Gemini 3.7 Flash as OpenAI Previews 14x-Faster GPT-5.6 Sol Ultrafast

Google Ships Gemini 3.7 Flash as OpenAI Previews 14x-Faster GPT-5.6 Sol Ultrafast

Author: Decrypt·

Key Takeaways

  • Gemini 3.7 Flash is generally available and can be used in more than 160 countries.
  • The model accepts up to one million input tokens and can process text, images, video, audio, PDFs, and tool calls.
  • Google says Gemini 3.7 Flash completed its test coding task in 2 minutes and 13 seconds, faster than the previous Flash model.
  • OpenAI previewed GPT-5.6 Sol Ultrafast as an invite-only API tier that uses Cerebras chips and reaches up to 750 output tokens per second.
  • Google priced Gemini 3.7 Flash at 75 cents per million input tokens and $3.75 per million output tokens through year-end before rates increase on December 31.
Google Ships Gemini 3.7 Flash as OpenAI Previews 14x-Faster GPT-5.6 Sol Ultrafast

Google launched Gemini 3.7 Flash on Thursday, a low-cost coding-and-agents model that is now generally available. On the same day, OpenAI previewed GPT-5.6 Sol Ultrafast, a Cerebras-powered service tier that runs its most capable model up to 14x faster. The speed race has shifted from raw intelligence to real-time agents — but only Google's model reaches every developer today.

Google and OpenAI both pushed the same message: AI is now fast enough to feel impressive for those who use AI agents, with each company announcing ultra-fast models. The two launches are built differently, and only one is actually in developers' hands today. The focus on speed is structural: agents can chain many model calls in sequence — planning, calling tools, checking results — so per-call latency compounds across a single task.

Google shipped Gemini 3.7 Flash, its latest model tuned for software coding and autonomous business workflows. OpenAI opened a limited preview of GPT-5.6 Sol Ultrafast, a new service tier that runs its most capable model at up to 750 output tokens per second.

Today we're introducing Gemini 3.7 Flash, our most intelligent workhorse model yet for coding and agents. This model brings substantial gains across software engineering, web development, and complex knowledge work. Now through the end of the year, Gemini 3.7 Flash is available… pic.twitter.com/RSCBDipjKn

— Google (@Google) August 13, 2026 (X post)

Gemini 3.7 Flash is a general-availability model. It accepts up to one million input tokens — roughly 750,000 words — and returns 64,000, handling text, images, video, audio and PDFs, and it can call tools and control a computer. Google is pitching it as the cheap brain for autonomous systems that plan tasks and finish multi-step jobs with less human help.

The model is not sacrificing quality for speed. It is both more capable and more efficient, completing Google's test coding task in 2 minutes and 13 seconds, whereas the latest Flash model took more than 5 minutes. The quality gap between the two is also noticeable.

OpenAI's Ultrafast isn't a new model. It is GPT-5.6 Sol — the same model OpenAI used an AI red team to harden against prompt-injection attacks before launch — on a faster track, powered by chipmaker Cerebras. It is around 14 times faster than GPT-5.6 Sol's own standard speed.

Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed. Launching first in the OpenAI API to a select group of customers with expanded access to more businesses as capacity grows. pic.twitter.com/a5dleofiDJ

— OpenAI (@OpenAI) August 13, 2026 (X post)

Cerebras' wafer-scale chips (processors fabricated at nearly the size of a full silicon wafer, rather than the many small chips usually cut from one) generate up to 750 tokens a second — about 560 words — fast enough that a voice agent can think mid-call.

The numbers that matter

Google's own benchmark sheet puts Gemini 3.7 Flash ahead of Claude Sonnet 5, GPT-5.6 Terra, and others on 11 of 18 tested categories, including a top Code Arena web-development score of 1,588 Elo and 30.4% on AutomationBench for enterprise workflows. That is all based on Google's methodology, so the lead should be treated as the company's claim.

As with any Gemini Flash model, it is also cheap. At 75 cents per million input tokens and $3.75 per million output tokens through year-end, it costs half of Gemini 3.6 Flash's original rate. That introductory price expires December 31, then doubles to $1.50 and $7.50 — which is still inexpensive for a Google model.

OpenAI hasn't published head-to-head scores for Ultrafast beyond customer quotes. Jane Street AI engineer John Crepezzi said in OpenAI's announcement that Cerebras' speed "enables different ways of using the models." Podium product lead Courtland Lykins called it "invaluable in our voice stack," saying the speed "completely changes the call experience." The tier is invite-only for now.

The speed push lands as the labs pivot from "who's smartest" to "who's fast enough for agents." Google's timing is pointed. Its flagship Gemini 3.5 Pro is still missing, with no release date given, three weeks after 3.6 Flash and days after a DeepMind leadership reshuffle that moved Demis Hassabis aside for deputy Koray Kavukcuoglu. OpenAI, meanwhile, is renting Cerebras' speed rather than waiting on its own stack.

Gemini 3.7 Flash is live now in more than 160 countries; GPT-5.6 Sol Ultrafast remains invite-only.