The Voice That Thinks: OpenAI Raises the Curtain on GPT-Realtime | OpenAI stock | Chat OpenAI | ChatGPT 4 | Turtles AI
OpenAI has officially made the Realtime API available with the gpt-realtime speech-to-speech model, which improves speech quality, instruction comprehension, and function calls, with support for images, SIP, and remote MCP servers, at a 20% lower cost than the preview.
Key points:
- Realtime API now generally available, no longer in beta.
- gpt-realtime, new speech model, more natural and precise.
- Extended features: visual input, SIP, remote MCP server.
- 20% reduced price on audio input/output tokens.
At the heart of the new development announced on August 28, 2025, the Realtime API is now out of beta, which began last October. This progress represents a milestone for developers: finally, a stable and straightforward interface for building voice agents ready for everyday use.
The star of the announcement is the gpt-realtime model, an all-in-one speech-to-speech solution: a single model processes incoming audio and outputs audio, dramatically reducing latency compared to traditional pipelines that combine STT, LLM, and TTS. The result is more streamlined, natural, and seamless conversations—almost like talking to a person.
Regarding speech, gpt-realtime delivers a more natural, expressive voice capable of following complex commands, such as reading a disclaimer word-for-word, repeating alphanumeric sequences, or seamlessly switching between languages. The two new voices, Cedar and Marin, are available only in the Realtime API, while the eight existing voices are receiving updates to benefit from it.
From a technical standpoint, gpt-realtime shows significant improvements on three internal benchmarks:
Big Bench Audio (audio reasoning): 82.8% vs. 65.6% for the December 2024 model.
MultiChallenge Audio (following instructions): 30.5% vs. 20.6%.
ComplexFuncBench (function calling): 66.5% vs. 49.7%.
Additionally, asynchronous tool calling is supported, allowing for seamless conversations even when the model is waiting for responses from external functions.
Regarding API features, we now have:
Support for remote MCP servers: simply specify a URL and the API automatically handles tool calls, simplifying integration.
Visual input: images, photos, or screenshots can be sent along with audio or text, allowing the model to "see" and respond to what the user shows.
SIP Support: The API can connect directly to the telephone network, PBX systems, and SIP endpoints.
Reusable prompts, similar to those of the Responses API, allow you to save prompt configurations for reuse across sessions.
Security is not overlooked: the API integrates active classifiers to detect problematic content, supports EU data residency, anti-impersonation policies, and transparent standards for reporting AI use.
Finally, pricing: 20% off the gpt-4o-realtime preview. Current costs: $32 per million audio input tokens ($0.40 per cached token) and $64 per million audio output tokens.
This was echoed in the comments on Reddit, where it was noted:
"Realtime API is now GA with gpt-realtime (their best speech-to-speech yet): lower latency … more natural voices … better at instructions + function calls … support for images, SIP … pricing dropped ~20% too."
OpenAI has provided a mature and accessible tool, ready for real-world use in customer support, voice assistance, and education, eliminating costs and technical barriers.


