OpenAI’s New Voice AI Models Improve Transcription Accuracy


openai plus free
Image credit: OpenAI

OpenAI voice AI models now target real-time transcription and completed audio files. ChatGPT desktop recently received a major Voice upgrade, but OpenAI is also expanding its audio tools for developers.

OpenAI introduces two transcription models

OpenAI announced two API models called GPT-Live-Transcribe and GPT-Transcribe.

GPT-Live-Transcribe focuses on low-latency, real-time transcription for live conversations. GPT-Transcribe handles completed recordings and asynchronous batch-processing workloads.

Both models support context-aware transcription. Developers can provide keywords, names, industry terminology, and expected languages to help the models understand a recording.

GPT-Live-Transcribe can also use earlier transcription turns as context during an ongoing conversation.

Context improves transcription accuracy

On OpenAI’s Context Aware ASR benchmark, additional context increased GPT-Live-Transcribe’s semantic accuracy from 38.5% to 44.6%.

GPT-Transcribe improved from 41.6% to 45.2% on the same benchmark.

OpenAI also reported the following transcription error rates on real-world audio:

  • GPT-Live-Transcribe: 9.60%
  • GPT-Realtime-Whisper: 11.65%
  • GPT-Transcribe: 8.98%
  • Whisper-1: 15.21%

Artificial Analysis independently measured a 3.31% word error rate for GPT-Transcribe. OpenAI charges $4.50 per 1,000 minutes of audio processed with the model.

In other AI news, Microsoft introduced the MAI-Cyber-1-Flash model for detecting security vulnerabilities. Anthropic also said it does not support banning open-weight AI models, although the company remains concerned about potential misuse.

Via Neowin

More about the topics: AI, OpenAI

Readers help support Windows Report. We may get a commission if you buy through our links. Tooltip Icon

Read our disclosure page to find out how can you help Windows Report sustain the editorial team. Read more

User forum

0 messages