Zach Anderson
Nov 07, 2024 15:59
An in depth comparability of Common-2 and OpenAI’s Whisper fashions reveals Common-2’s superior efficiency in accuracy, correct noun detection, and decreased hallucination charges.
In a complete evaluation of main Speech-to-Textual content fashions, AssemblyAI’s Common-2 has emerged as a prime performer when in comparison with OpenAI’s Whisper variants, in accordance with a current report by AssemblyAI. The analysis targeted on real-world use instances, assessing fashions on duties important for creating correct transcripts, corresponding to correct noun recognition, alphanumeric transcription, and textual content formatting.
Mannequin Comparability
The evaluation in contrast Common-2 and its predecessor Common-1 with OpenAI’s Whisper large-v3 and Whisper turbo fashions. Every mannequin was evaluated primarily based on parameters like Phrase Error Price (WER), Correct Noun Error Price (PNER), and different metrics crucial for Speech-to-Textual content duties.
Efficiency Metrics
Common-2 achieved the bottom Phrase Error Price (WER) at 6.68%, marking a 3% enchancment over Common-1. Whisper fashions, whereas aggressive, had barely greater error charges, with large-v3 recording a WER of seven.88% and turbo at 7.75%.
In correct noun recognition, Common-2 demonstrated superior accuracy with a 13.87% PNER, outperforming each Whisper large-v3 and turbo. This mannequin additionally excelled in textual content formatting, attaining a U-WER of 10.04%, which signifies higher dealing with of punctuation and capitalization.
Alphanumeric and Hallucination Charges
Whisper large-v3 confirmed power in alphanumeric transcription with the bottom error price of three.84%, barely forward of Common-2’s 4.00%. Nonetheless, Common-2’s decreased hallucination charges have been a major benefit, with a 30% discount in comparison with Whisper fashions, making it extra dependable for real-world functions.
Conclusion
Common-2’s developments over Common-1 are evident, with enhancements in accuracy, correct noun dealing with, and formatting. Regardless of Whisper’s strengths in sure areas, its susceptibility to hallucinations poses challenges for constant efficiency.
For additional insights and detailed metrics, the complete analysis is out there by AssemblyAI’s official report.
Picture supply: Shutterstock


