API

Voicemail detection API: one WebSocket, one verdict, about two seconds

A real-time voicemail detection API for any dialer or voice platform. Stream the first seconds of the callee's audio over a WebSocket and get back HUMAN or MACHINE with a reason, before an agent or a bot says a word.

The API is built for the moment of answer. There is no upload, no job id and no callback: the client opens a connection when the call is answered, streams audio as it arrives, and reads a single JSON verdict. Most integrations stop streaming the moment the verdict arrives, about two seconds in.

The protocol in four steps

  • Open a WebSocket to ws://app.vmhunter.com:2701. One connection per call.
  • Send one text frame: {"config": {"vid": "<your call id>", "api_key": "<key>", "sample_rate": 8000, "bytes_per_sample": 2}}
  • Send the callee's audio as binary frames: 16-bit PCM, mono, 8 kHz, 20 ms per frame.
  • Read one text frame: {"AMDSTATUS": "MACHINE", "AMDCAUSE": "MACHINE_BEEP"}. The server closes the connection.
import json, time
from websocket import create_connection

ws = create_connection("ws://app.vmhunter.com:2701", timeout=10)
ws.send(json.dumps({"config": {"vid": "lead-1", "api_key": API_KEY,
                               "sample_rate": 8000, "bytes_per_sample": 2}}))
for chunk in audio_frames():        # 320 bytes = 20 ms of 8 kHz s16le
    ws.send_binary(chunk)
print(json.loads(ws.recv()))        # {"AMDSTATUS": ..., "AMDCAUSE": ...}

The full reference, with every AMDCAUSE value and the timing rules, is at /docs/api. Minimal Python and Node.js clients are in vmhunter-examples.

What you get back

AMDSTATUSAMDCAUSEMeaning
HUMANHUMANA live person: a greeting followed by a pause, a question, or a live-person phrase
MACHINEMACHINEVoicemail or IVR wording, or a greeting that ran on without pausing
MACHINEMACHINE_BEEP / MACHINE_STATICA voicemail beep; only noise on the line
MACHINEDISCONNECTSIT tones: the number is disconnected or changed
MACHINEINITIALSILENCENo speech in the window, including calls that arrive with no audio
MACHINE or CALLGUARDCALLGUARD_PHRASE:<phrase>Call screening (iPhone, Google, call blockers); the matched phrase follows
FAILEDAUTH_FAILED, LIMIT_REACHED, ...Detection did not run; not billed

What it is built for

  • Dialers that need to connect or drop a call before the agent hears anything: VICIdial, Asterisk and FreeSWITCH already have bundled clients.
  • Voice AI agents that must not start their script to a voicemail greeting. See voicemail detection for AI voice agents.
  • CPaaS-built apps on Twilio Media Streams or similar, where you control the audio stream and the call.
  • Anyone who wants the evidence. Every call is logged with its transcript, reason and audio sample, and exportable as CSV.

Limits and pricing

Plans are priced per call analysed and set the calls per second and concurrent connections you can open. The free plan includes 5,000 calls a month with no card; paid plans start at $100 a month. Details on the pricing page.

FAQ

Questions

What does the voicemail detection API return?

One JSON frame per call: AMDSTATUS (HUMAN, MACHINE, CALLGUARD or FAILED) and AMDCAUSE, the reason, such as MACHINE_BEEP, DISCONNECT, INITIALSILENCE or CALLGUARD_PHRASE with the matched phrase. The server then closes the socket.

Is it REST or WebSocket?

WebSocket. Detection needs the audio as it happens, so the client opens one connection per call, sends a config frame with the API key, streams raw PCM, and reads the verdict. There is no polling and no callback to host.

What audio format is required?

Signed 16-bit little-endian PCM, mono, 8 kHz, no headers, which is what telephony delivers. Send it in 20 ms frames as it arrives; the engine classifies as the audio streams in.

How fast is it?

The engine decides on roughly the first two seconds of audio and listens up to three seconds if the callee is still talking. A verdict typically arrives about 2.3 seconds after the first frame. Anything you buffer before streaming delays it by the same amount.

What happens on an error?

AMDSTATUS=FAILED with a cause: AUTH_FAILED, LIMIT_REACHED, CONFIG_TIMEOUT, or SPEECHLLM_ENGINE_UNAVAILABLE if the engine could not process the call. FAILED calls are not billed. Clients should apply their own fail-safe, usually MACHINE.

Are there SDKs?

The protocol is simple enough that a client is a few dozen lines. Minimal Python and Node.js examples are on GitHub, and the Asterisk, FreeSWITCH and Twilio clients are open source if you want a fuller reference.

Get a key and send a call

The free plan includes 5,000 calls a month. The Python example above runs in under a minute.