archive
· today in ai · 2026-07-15
Study finds voice AI models don't generalize
Archive item — written before sources were shown.
Hume AI's Real World VoiceEQ benchmark rated 40-plus voice models on ~1M human judgments and found each one specializes rather than excelling across the board.
The benchmark combined roughly 785,000 text-to-speech and 48,000 speech-to-speech human ratings across 15-plus dimensions and 60-plus metrics, and found most models miss paralinguistic cues like tone and hesitation even when they score well on raw intelligibility.
sources
- 01Introducing Real World VoiceEQ: Measuring the human quality of voice AIhuggingface.co · primary announcement
