
OpenAI Realtime API now supports voice agents with sub-300ms latency
Quick Answer
OpenAI's Realtime API now enables voice agents with sub-300ms first-token latency, enhancing user interaction with features like barge-in handling and on-the-fly memory updates.
Quick Take
Additionally, pricing for cached prompts has been reduced by 30%, making it more cost-effective for developers.
Key Points
- Sub-300ms first-token latency improves responsiveness for voice agents.
- Features include barge-in handling and on-the-fly memory updates.
- Cached prompt pricing has decreased by 30%, benefiting developers.
- Enhanced user experience through faster interaction capabilities.
- Real-time applications can leverage these improvements for better performance.
Article Excerpt
From source RSS / original summaryThe OpenAI Realtime API now supports tool-using voice agents with sub-300ms first-token latency, including barge-in handling and on-the-fly memory updates. Pricing drops 30% for cached prompts.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from OpenAI Blog
See more →Scientific computing in the age of agentic AI
AI agents are transforming scientific computing by streamlining software development, enabling researchers to focus on discovery. Projects using Codex and Claude Code report accelerated development and improved maintenance, though challenges in validating AI outputs remain. Long-term stewardship of research software is crucial to ensure reliability and reproducibility.