Engineering

Acoustic AMD vs Speech-Based AMD: We Built an Acoustic Model and Tested Both on 4,000 Agent-Labelled Calls

Is answering machine detection better done on sound alone or on the words? We built an acoustic model from 500,000 recordings and ran it against VM Hunter on 4,000 calls labelled by agents. The results, the method, where acoustic is the better tool, and the limits of the test.

Marketing Team

VM Hunter

October 4, 2026
8 min read
Acoustic AMD vs Speech-Based AMD: We Built an Acoustic Model and Tested Both on 4,000 Agent-Labelled Calls

There are two ways to build answering machine detection with machine learning. An acoustic detector classifies the sound of the first seconds of a call and never looks at words. A speech-based detector, which is what VM Hunter is, hears the words, reads them, and also measures the audio.

You will find confident claims online that one approach is simply better than the other. We prefer to measure. So we built an acoustic model ourselves, from a dataset of more than 500,000 call recordings, and ran it against VM Hunter on 4,000 calls that human agents had labelled. This post is the result, with the method and the limits.

The short version

On calls where the first word is a lone "Hello", the hardest case in outbound dialing:

Aligned test set (2,271 calls)Answering machines passed to an agent (lower is better)Live people passed to an agent (higher is better)
VM Hunter (speech-based)62.3%93.7%
Acoustic model, full size78.7%95.4%
Acoustic model, lite74.2%93.8%
  • VM Hunter caught 313 of 831 answering machines. The acoustic model caught 177.
  • Of those 177, only 15 were machines VM Hunter missed.
  • Most of the live calls VM Hunter appears to lose here are mislabelled recordings, not people. More on that below.

The acoustic model we built

It is not a toy. It is a convolutional network on log-mel spectrograms plus a DSP tone detector, built from a dataset of 518,693 unique two-second recordings (trained on the reliably labelled part) in five classes: human, recorded speech, call screening, silence and tone. It decides at 2.0 seconds and serves 5,000 concurrent calls on one machine.

On its own held-out test set it scores a macro-F1 of 0.977, with about 2.2% of live people lost and 1.1% of recordings passed as human. By the usual standards, that is a finished model.

One caution: this is our acoustic model, not anyone else's product. Another vendor's acoustic detector may do better or worse. What the test shows is what the sound of the first two seconds can and cannot tell you, and that part applies to any model that only listens to sound.

The test data

The calls come from one dialer deployment over one week in September 2026. For each call there is the recording the detector judged and the disposition the agent chose after taking the call. That is the best ground truth a dialer has: a person listened and said what it was.

We took the 4,000 calls where the callee's first word was a lone "Hello" and an older build of our engine had passed the call to an agent:

What the agent foundCalls
Answering machine1,500
Live conversation1,500
No answer after connect600
Dead air or disconnected400

Neither model was trained on these calls. And note what the set is: every call here already fooled a detector once. Easy voicemail greetings, beeps and carrier prompts are not in it.

How each model was run

Acoustic model. Exactly as it runs in production: the first 2.0 seconds of audio into the network and the tone detector, human threshold 0.4.

VM Hunter. We replayed the production engine's rules and audio functions offline. VM Hunter listens for 2.0 seconds, and if the callee is still talking and the words so far are not decisive, it keeps listening until a 300 ms pause or 3.0 seconds. The live speech engine cannot be called offline, so for the calls where VM Hunter would have listened longer (39% of them), a second, independent speech model transcribed exactly the audio VM Hunter would have heard. That makes the VM Hunter figure an estimate, and we say so in the limits.

Aligned subset. These recordings start after a played prompt, so on many of them speech begins late. An acoustic model that sees only 2.0 seconds would be judged on silence. To be fair to it, the headline table uses the 2,271 calls where speech starts within the first second.

Results

Aligned subset: percentage passed as HUMAN

Machine (831)Live (867)No answer (347)Dead air (226)
Older engine build100100100100
VM Hunter62.393.783.689.4
Acoustic, full78.795.489.993.8
Acoustic, lite74.293.890.592.5

Which model caught which machines (of 831)

BothOnly VM HunterOnly acousticNeither
16215115503

Can the acoustic model be tuned to match? Its human threshold trades machines for people:

ThresholdMachines passedLive people lost
0.4 (default)78.7%4.6%
0.673.0%6.5%
0.864.7%9.7%

Even at 0.8 it passes more machines than VM Hunter (62.3%) while losing 9.7% of live people against VM Hunter's 6.3%.

All 4,000 calls, without the alignment filter: VM Hunter passed 69.2% of machines and 94.7% of live people; the acoustic model passed 81.9% and 89.1%.

Why the words win

A voicemail greeting that starts "Hello" and a person who says "Hello?" make the same sound for the first second. The difference is in what follows:

  • "Hello?" and then a pause: a person.
  • "Hello, we are not available now": a recording.
  • "Hello, please state your name after the tone": call screening.

An acoustic model has to infer that from rhythm and voice quality. A speech-based detector reads it. When the callee is still talking at two seconds, VM Hunter listens a little longer, and the next three or four words usually settle it.

We also measured whether an acoustic model can decide earlier than two seconds. At 1.0 second, our model lost 21% of live people: a person and a recording have not yet diverged. A detector can start listening in milliseconds, but the evidence for a human is the pause after the greeting, and that takes time to arrive whichever method you use.

The live calls VM Hunter "lost"

Across all 4,000 calls, VM Hunter marked 80 of the 1,500 agent-labelled live calls as machine. We read what was heard on every one. About 70 contain voicemail or screening wording, such as "Hello, we are not available now" (27 calls) and "calls to this number are being screened". The agent had dispositioned them as not interested or call back, but the audio is a recording. Real misses were about 10, or 0.7% of live calls.

This is another thing words give you: every VM Hunter verdict comes with the transcript and the reason, so a disputed call can be checked by reading it.

What VM Hunter does with sound

Speech-based does not mean words only. VM Hunter analyses the audio on every call as well:

  • A beep detector and a SIT tone detector for voicemail beeps and disconnected numbers.
  • Silence and no-audio detection, reported with its own cause so a dialer or carrier fault can be told apart from a quiet line.
  • Static and tone detection for ringback, carrier pips and noise.
  • Audio tail analysis: whether the callee is still talking at the cutoff, which is what triggers listening longer.

So tones, beeps and dead air are classified from the audio, with no words needed. VM Hunter's speech layer is our SpeechLLM Engine, a streaming speech engine; it is not Whisper.

Where acoustic is the better tool

An honest comparison includes this:

  • Cost and footprint. No speech engine. Our lite model serves 5,000 concurrent calls on CPU.
  • Language. An acoustic model does not depend on vocabulary. VM Hunter is tuned for English. (Our acoustic model was trained on English-language calls, so we have not measured this.)
  • The easy majority. On long greetings, carrier prompts, silence and tones, the acoustic model is accurate. Its weakness is the short answer.

And one case neither approach solves: a person who answers and says nothing for two seconds sounds exactly like dead air. Our acoustic model calls it silence; VM Hunter reports it as initial silence. Both tell the dialer what was heard and let it decide.

Limits of this test

  • It is the hardest slice, not the general mix. Every call is a lone "Hello" that fooled an older build. On ordinary traffic, where most machines are long greetings, both do far better.
  • The VM Hunter figure is an estimate. A stand-in speech model replaced the live one for the listen-longer calls. If the live engine heard no more than the first "Hello", VM Hunter would have passed these calls as the older build did.
  • Agent labels contain errors, as the 70 calls above show.
  • One site, one week, English-language calls.
  • 61% of these machines were caught by neither model. One word and then silence is, in two to three seconds, the same as a person. That is the open problem for everyone, and anyone who tells you otherwise should show their test.

Run this on your own calls

You do not need to take our numbers. Pull a few hundred calls your agents dispositioned as answering machine and a few hundred they had real conversations on, and replay the first three seconds through any detector you are considering. Count the two errors separately. The method is in How We Test VM Hunter's Accuracy.

VM Hunter works with VICIdial, Asterisk, FreeSWITCH and Twilio, and every verdict is logged with its transcript, reason and audio in the call logs.