News

OpenAI introduces two transcription models: low latency and better understanding of context

1 min read
Source: ithome.com
IT House reported on July 29 that OpenAI company issued an announcement on the In terms of calling costs, the GPT-Live-Transcribe transcription fee is US$0.017 per minute (IT House Note: the current exchange rate is approximately 0.12 yuan). The GPT-Transcribe transcription fee is US$0.0045 per minute (the current exchange rate is approximately 0.03 yuan) OpenAI Officials say the two transcription models can better understand context and provide more accurate transcription on real-world audio in a variety of accents and languages than the company's previous models, including phrases, numbers, jargon, and speech with loud background noise. GPT-Live-Transcribe is built for low-latency real-time transcription; GPT-Transcribe optimizes asynchronous transcription of completed audio files and batch workloads. On the Context Aware ASR benchmark, GPT-Transcribe's semantic accuracy improves from 41.6% without free-form context to 45.2% with context. Across Common Voice's 22 languages, GPT-Transcribe has a transcription error rate of 19.27%, while Whisper (whisper-1) has a transcription error rate of 40.37%. On the real-world audio recording benchmark across nine languages, it achieved an 8.98% transcription error rate, compared to 15.21% for Whisper (whisper-1).