How We Test VM Hunter's Accuracy: Method, Sample Sizes and Where It Misses
Every AMD vendor publishes an accuracy number; almost none publish the method. Here is ours: replay against 78,000 transcripts, 7,000 recordings checked by verdict type, a second speech model as a cross-check, the results by category, and the five cases the engine still gets wrong.
Marketing Team
VM Hunter

Every answering machine detection vendor publishes an accuracy number. Almost none publish how it was measured, on what, or where the detector fails. That makes the numbers hard to compare and easy to doubt, and customers are right to ask.
This post is our answer to that question. It describes how VM Hunter is tested, with the sample sizes, what the tests found, and the cases the engine still gets wrong. The method is the part you can reuse on any detector, including ours.
Three kinds of test
We use three, because each catches something the others miss.
1. Replay against transcripts. Every call VM Hunter analyses is logged with the transcript the engine heard and the verdict it gave. When a rule changes, the new rules are replayed over the stored transcripts and every verdict that changes is listed. Before the September 2026 engine release, that replay ran over about 78,000 transcripts from our own test campaigns and 1,269 from a customer's labelled set. The output is not an accuracy number; it is the list of calls whose verdict would flip, each read by a person. A change ships only when every flip is in the right direction.
2. Listening to recordings by verdict. The engine keeps a short audio sample of what it heard. For each test run we pull recordings by verdict type, measure them (level, duration of speech, tone content), and listen to the ones that matter. Rare verdicts are checked in full; common ones are sampled. Across the pre-release runs that covered about 7,000 recordings: every silence and tone verdict, every "empty transcript", every HUMAN, and samples of a few hundred from the voicemail-greeting class.
3. Second-opinion transcription. For recordings where the question is "was there speech the engine missed?", we run a second, independent speech model over the audio and compare. It is slower than the live engine and not usable for real-time detection, but it is a good check on the live engine's hearing.
What we score, and what we do not
A verdict is scored against what was audible in the detection window, the first two to three seconds after answer. Not against how the call ended.
The distinction matters more than anything else in this post. A call where a screening prompt plays and a person picks up at twelve seconds was, at two seconds, a screening prompt. A call where nothing is audible for the first two seconds cannot be classified by any detector, and scoring it as a miss measures the trunk, not the engine. One customer's 36 to 40% "accuracy" figure became 96% when their own files were re-scored this way; the difference was entirely in what was being measured.
We count two kinds of error separately, because they cost different things:
- False MACHINE: a live person called a machine. The dialer hangs up on a reachable lead.
- False HUMAN: a machine called a person. An agent sits through a greeting.
What the pre-release tests found
The September 2026 release was tested on about 94,000 calls over three days on a test trunk, at peaks of 70 calls per second, plus the customer-labelled sets. The list was deliberately hard: about 80% of answered calls were carrier voicemail prompts, and only 1.5 to 2% were live people. That mix is good for stress and bad for a headline number, so here are the component results instead.
| Verdict class | Recordings checked | Result |
|---|---|---|
| No audio (socket open, zero or silent bytes) | 334 | All but 5 were pure digital silence; the rest held a few bytes of noise. Correct. |
| No speech (line hiss only) | 651 | Peaked around -66 dBFS, line hiss only. Correct. |
| Static / tone | 559 | All the same carrier pip or ringback. Correct. |
| Voicemail greeting matched by wording | 900 sampled | Continuous speech in every sample. Correct. |
| Empty transcript (audible, nothing transcribed) | 863 | Mostly tones and clicks. About 40 were Spanish voicemail (correct verdict). About 8 in 46,000 calls were a real "Hello" the speech engine did not catch. |
| HUMAN | about 1,350 (all of them) | About 95% correct or defensible. The wrong ones were carrier prompts that paused right at the cutoff; each wording was added to the rules. |
The customer-labelled set gave a second view: on the 1,269 recordings that customer's dialer had passed as HUMAN, the audio proved 162 were answering machines (the voicemail wording arrived in the third second). After the rule changes, 24 of those are caught from the first two seconds alone and a further 93 now trigger extended listening.
We do not publish a single number from these tests, because the list mix decides it. On a dead list the "accuracy" is dominated by voicemail, which is the easy case; on a fresh consumer list it is dominated by short "Hello?" answers, which is the hard one.
Where it still misses
These are the open cases as of October 2026. If your traffic is heavy in one of them, test before you trust.
- Name-only answers. "Hi, this is Kim." followed by a pause is how a person answers and also how a personal voicemail greeting starts. The engine calls it HUMAN after a short pause. Some of those are voicemail. Flipping them would drop real people, so they stay HUMAN; an optional longer-pause check exists for customers who prefer the other trade.
- Greetings that pause exactly at the cutoff. A recording that pauses at two seconds looks like a person who stopped talking. Extended listening catches most of these; a few get through.
- Non-English greetings. The engine is tuned for English. A Spanish carrier prompt produces an empty transcript, which is classified MACHINE by the audio shape. That is the right verdict for voicemail but means a Spanish-speaking person is not recognised by their words.
- Late pickups behind a prompt. A voicemail prompt starts, then a person picks up within the window ("Please leave... Hello?"). The voicemail wording wins. This is rare in our logs (single digits in tens of thousands of calls) but real.
- The speech engine's hearing. A very short or very quiet "Hello" is occasionally not transcribed at all. Measured at roughly 1 in 5,000 calls on the test list.
How to run this on your own calls
You do not need our tools to do this. You need recordings and twenty minutes.
- Pull 200 calls at random from one campaign day, 100 the detector called HUMAN and 100 it called MACHINE.
- Listen to the first three seconds of each. Mark what was actually there: person, voicemail, carrier prompt, call screening, silence.
- Count false MACHINE and false HUMAN separately.
- Ignore the calls with nothing audible in the window when scoring the detector; count them separately as a trunk number.
If a vendor's number does not survive that, the number was measuring something else. If ours does not survive it on your list, tell us, and send the recordings: that is how the rules improve.
VM Hunter logs the transcript, the reason and the audio for every call, which turns step 2 from listening into reading for most of the sample. The call logs filter by verdict and cause, and export to CSV.