Almost Human, Except When It Matters: VoxParity and the Decisions a Voice Should Change
Abstract
A voice agent can handle almost every call on the words alone and still fail the few its sector's rules were written for. Emergency-call standards, fraud guidance, radio phraseology and vulnerability rules recognise that how a caller sounds, or what else is audible, can change the right action. VoxParity tests whether agents act on it. In 183 scenarios from 14 sectors, one transcript stays fixed while the audio changes (a coaching voice, a medical monitor beeping, a mayday under a radio check, noise over a drug name, a child's voice placing a bet, a frightened whisper), and with it the correct typed tool call. A words-only null test credits a system only if hearing the call moves its actions more than it moves a pipeline that only reads the words. Only 11 of the 23 systems that can also be run on the transcript pass. Descriptively, errors run toward the words: when the audio calls for protection, all 28 systems carry out the routine request more often than they over-react on clean calls (41% against 12% pooled; the words-only pipeline, 58% against 15%). Exploratory analyses place most of the leading systems' misses on cues they heard; systems beat the null almost entirely on items that state the rule; the leading systems overrule heard resignation or confusion far more often than acute alarm; and, in the models tested, describing the voice and stating the rule each recover part of the shortfall, leaving a gap on emotion.
Community
A caller agrees to a power disconnection while a medical monitor beeps in the background. All four leading systems we tested identify the monitor in their own probe; three schedule the disconnection, and none of the 28 systems applies the vulnerability hold that utility rules require.
VoxParity measures whether voice agents act on what the audio tells them. In 183 scenarios from 14 sectors, the transcript stays fixed while the audio changes (emotional delivery, a second voice, a background sound, disfluency or silence, speaker age, sarcasm, a masked word), and with it the correct typed tool call. Scoring is deterministic on the executed call, with no LLM judge. We test 28 systems from 11 vendors, including 9 production realtime agents, with about 20 volunteers as a human reference.
What we find:
- Errors run toward the words. When the audio calls for protection, all 28 systems carry out the routine request more often than they over-react on clean calls (41% vs 12% pooled; a words-only pipeline 58% vs 15%; descriptive).
- Half of the systems we could test show no measurable use of the audio. Their actions can't be told apart from a pipeline that never hears the call (12 of the 23 systems with a transcript path, including 6 of the 7 realtime agents).
- Facts move agents more than feelings. Given a one-word label of the cue, gemini-3.7-flash acts correctly on 1.00 of background-sound calls and 0.97 of second-voice calls, but 0.59 of emotional ones (exploratory; an upper bound).
- The frontier's loss is in deciding. For the four leading systems, perfect hearing would add 0.04 credit and perfect deciding 0.28 (exploratory).
Resources:
- Code, judge-free scorer and bring-your-own-agent harness: https://github.com/bhavik-mangla/voxparity-bench
- Development split, 40 scenarios with audio (CC BY 4.0): https://huggingface.co/datasets/bhavikmangla/voxparity-dev
- Leaderboard with 95% intervals: https://bhavik-mangla.github.io/voxparity-bench/
The other 143 scenarios are held out with published hashes; evaluation on them is available on request.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents (2026)
- DuplexWorld: Can voice agents help you get through the day? (2026)
- SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning (2026)
- The Machines Are Calling: Measuring Automated and Synthetic Voices in Unwanted Inbound Calls (2026)
- SEABED: SouthEast Asian Benchmark for Evaluating Audio Reasoning (2026)
- Inquesto Score: A reliability Protocol For Voice Agents (2026)
- Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.35922 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper