The release shows that the focus of the AI race is shifting from raw performance to the speed of running real-time agents. [Photo: ChatGPT]

[DigitalToday reporter Jinju Hong] Google and OpenAI have rolled out ultrafast artificial intelligence services aimed at coding and the AI agent market. The AI race appears to be shifting beyond model performance to how quickly results can be produced.

On Aug. 13 (local time), blockchain media outlet Decrypt reported that Google officially launched the low-cost model "Gemini 3.7 Flash," tailored for coding and AI agent tasks. On the same day, OpenAI unveiled "GPT-5.6 Sol Ultrafast," which significantly boosts the speed of its existing GPT-5.6 Sol.

The two companies differ in how they are rolling the products out. Gemini 3.7 Flash is available immediately in more than 160 countries, but GPT-5.6 Sol Ultrafast is currently offered on a limited basis only to some API customers.

Google positioned Gemini 3.7 Flash as a flagship model suited to software engineering, web development and complex knowledge work. It especially targets being a "low-cost brain" for AI agents that plan and execute multiple steps on their own. It supports up to 1 million input tokens and 64,000 output tokens, and can process not only text but also images, video, audio and PDFs. It also provides tool-calling and computer-control functions.

Speed improvements are also central. According to Google, Gemini 3.7 Flash completed a test coding task in 2 minutes 13 seconds. That is a sharp reduction from the previous Flash model, which took more than 5 minutes.

Google also set pricing aggressively. It will offer the model through the end of this year at $0.75 per 1 million input tokens and $3.75 per 1 million output tokens. That is about half the initial price of Gemini 3.6 Flash. From next year, the prices rise to $1.50 and $7.50, respectively.

Google claimed in its own benchmarks that Gemini 3.7 Flash outperformed Claude Sonnet 5 and GPT-5.6 Terra in 11 of 18 evaluation categories. It recorded 1,588 Elo on Code Arena and 30.4 percent on AutomationBench. However, those figures should be considered as results based on Google's own evaluation methodology.

OpenAI on the same day introduced "GPT-5.6 Sol Ultrafast" as a new high-speed processing tier for GPT-5.6 Sol. It is not a new model release but a service that processes the existing GPT-5.6 Sol at ultrafast speed. OpenAI said it can respond up to 14 times faster than the standard speed by using chips from semiconductor company Cerebras.

The maximum output speed is 750 tokens per second. It is expected to be particularly useful for services that must answer immediately while conversing with users, such as real-time voice agents.

John Crepege, an AI engineer at Jane Street, assessed that Cerebras' processing speed could change how AI models are used. Cortland Likins, product head at Podium, also said fast response times in voice services could significantly change the calling experience.

Accessibility still lags behind Google's. GPT-5.6 Sol Ultrafast is currently available in the OpenAI API only for some customers, and OpenAI plans to expand the eligible companies as processing capacity increases.

The rivalry shows that as AI agents spread, response speed is emerging as a new competitive metric. AI agents plan multi-step tasks, call tools and produce results after receiving a user's request. In that process, not only the model's reasoning ability but also how much waiting time is reduced at each step directly affects the actual user experience. The importance of speed is even greater for voice-based AI, because the longer the response is delayed, the less it feels like a conversation with a person.

Google has put out a model that developers and companies can use immediately, emphasizing low prices and broad access. OpenAI, by contrast, chose a strategy of using Cerebras' computing infrastructure to push processing speeds of a top-tier model to an extreme.

Ultimately, this competition is expanding beyond which AI is smarter to which AI can handle real work faster and more cheaply.

In this situation, the next moves by the two companies are also drawing attention. A release schedule has not yet been disclosed for Google's flagship high-performance model, Gemini 3.5 Pro. OpenAI has moved to respond to the market by using Cerebras' processing speed instead of its own infrastructure. As ultrafast responses emerge as a core competitive strength for AI agents, the center of competition is expected to shift not only to model intelligence but also to real-world service speed and delivery methods.

Keyword

#Google #OpenAI #Gemini 3.7 Flash #GPT-5.6 Sol Ultrafast #Cerebras
Copyright © DigitalToday. All rights reserved. Unauthorized reproduction and redistribution are prohibited.