brief researchproducts
Study finds voice AI models don't generalize
Hume AI's Real World VoiceEQ benchmark rated 40-plus voice models on ~1M human judgments and found each one specializes rather than excelling across the board.
The benchmark combined roughly 785,000 text-to-speech and 48,000 speech-to-speech human ratings across 15-plus dimensions and 60-plus metrics, and found most models miss paralinguistic cues like tone and hesitation even when they score well on raw intelligibility.
sources 1 cited
1 huggingface.co Introducing Real World VoiceEQ: Measuring the human quality of voice AI