OpenAI’s New Voice AI Models Improve Transcription Accuracy
OpenAI voice AI models now target real-time transcription and completed audio files. ChatGPT desktop recently received a major Voice upgrade, but OpenAI is also expanding its audio tools for developers.
OpenAI introduces two transcription models
OpenAI announced two API models called GPT-Live-Transcribe and GPT-Transcribe.
GPT-Live-Transcribe focuses on low-latency, real-time transcription for live conversations. GPT-Transcribe handles completed recordings and asynchronous batch-processing workloads.
Both models support context-aware transcription. Developers can provide keywords, names, industry terminology, and expected languages to help the models understand a recording.
GPT-Live-Transcribe can also use earlier transcription turns as context during an ongoing conversation.
Context improves transcription accuracy
On OpenAI’s Context Aware ASR benchmark, additional context increased GPT-Live-Transcribe’s semantic accuracy from 38.5% to 44.6%.
GPT-Transcribe improved from 41.6% to 45.2% on the same benchmark.
OpenAI also reported the following transcription error rates on real-world audio:
- GPT-Live-Transcribe: 9.60%
- GPT-Realtime-Whisper: 11.65%
- GPT-Transcribe: 8.98%
- Whisper-1: 15.21%
Artificial Analysis independently measured a 3.31% word error rate for GPT-Transcribe. OpenAI charges $4.50 per 1,000 minutes of audio processed with the model.
In other AI news, Microsoft introduced the MAI-Cyber-1-Flash model for detecting security vulnerabilities. Anthropic also said it does not support banning open-weight AI models, although the company remains concerned about potential misuse.
Via Neowin
Read our disclosure page to find out how can you help Windows Report sustain the editorial team. Read more
User forum
0 messages