Skip to content

Benchmarks

AMD performance. Measured honestly.

Our benchmark methodology, the metrics we track, the languages we test, and the limitations we acknowledge. Where we do not have a figure measured on traffic like yours, we say so rather than quoting one.

Quick answer

What do the DialerBee benchmarks measure? They measure AI AMD accuracy: the model classifies a real-time transcript of the first 3 seconds of call audio as a live human or an answering machine, and we benchmark that across 11 languages, carrier environments and campaign types. Ground truth comes from agent overrides, where an agent who disagrees with a classification flags it in one click, and evaluation is continuous rather than a one-time lab test. Five metrics are tracked: accuracy, false positive rate, false negative rate, detection latency, and agent override rate. Published figures come from internal pilots and hold in selected pilot conditions; where we do not have a figure measured on traffic like yours, we say so rather than quoting one.

What We Benchmark

AMD accuracy across languages

DialerBee's AI AMD uses transcript classification on the first 3 seconds of call audio to determine whether a human or an answering machine has picked up. We benchmark this system across multiple languages, carrier environments, and campaign types to measure real-world accuracy.

Unlike vendors who report lab-condition accuracy, our benchmarks reflect production data with all the noise, carrier variability, and regional differences that come with real outbound campaigns.

Methodology

How we measure accuracy

Transcript Classification

The AMD model classifies a real-time transcript of the first 3 seconds of audio. It reads words and phrases, not audio waveforms or beep patterns.

Agent Override as Ground Truth

When agents disagree with the AMD classification, they flag it with one click. These overrides serve as ground truth labels for measuring accuracy.

Continuous Evaluation

Benchmarks are not one-time tests. We continuously evaluate accuracy against the stream of agent overrides across all production campaigns.

Metrics Tracked

Five metrics that matter

Accuracy

Overall correct classification rate across all calls.

False Positive Rate

Human calls incorrectly classified as machines.

False Negative Rate

Machine calls incorrectly classified as human.

Detection Latency

Time from call pickup to AMD classification decision.

Agent Override Rate

Percentage of calls where agents correct the AMD decision.

Benchmark Summary

Public summary from internal pilots

Detailed per-language, per-carrier and per-region figures from our internal pilots are available on request.

Metric Public Summary Notes
Calls evaluated 10,000–50,000 range Across internal pilot campaigns
Date range Q1–Q2 2026 Updated as new pilot data becomes available
Languages tested 9 EN, AR, ES, FR, IT, DE, TR, HI, UR
Regions tested MENA, Europe, North America Carrier mix varies by region
Campaign types Collections, BPO, sales, insurance Results segmented where possible
AMD false positive rate Below 3% in selected pilots Human incorrectly classified as machine
AMD false negative rate Below 5% in selected pilots Machine incorrectly classified as human
Median detection latency Sub-1 second Pickup to AMD decision
P90 detection latency Under 1.5 seconds Pickup to AMD decision
Agent override rate 2–5% range Manual correction rate in pilot campaigns
Contact-rate lift Up to 4x vs manual dialing In pilot workflows; results vary
Idle-time reduction 30–45% range In selected pilot campaigns

Based on internal pilot benchmarks. Results vary by campaign type, list quality, carrier, region, language, pacing configuration, and agent workflow. Per-language and per-carrier behaviour varies enough that we would rather measure it on your routes during a pilot than publish an average that will not match what you see.

Languages Tested

Eleven languages, real pilot data

EN

English

AR

Arabic

ES

Spanish

FR

French

IT

Italian

DE

German

TR

Turkish

HI

Hindi

UR

Urdu

Arabic benchmarks include Gulf, Levantine, and Egyptian dialect variants. Hindi and Urdu benchmarks include regional variant coverage. Each language is benchmarked independently with carrier-specific and region-specific breakdowns.

Limitations

What benchmarks cannot tell you

Results vary by carrier. Different carriers produce different voicemail greetings, audio quality, and connection timing.

Results vary by campaign type. Collections campaigns have different pickup patterns than sales or survey campaigns.

Results vary by region. Voicemail greeting styles, carrier infrastructure, and network conditions differ across markets.

Results vary by list quality. Stale numbers, wrong numbers, and disconnected lines affect AMD performance differently.

Benchmarks reflect aggregate performance. Individual campaign results may be higher or lower than reported averages.

Full internal-pilot benchmark data on request

Internal-pilot benchmark reports with per-language accuracy tables, carrier-specific breakdowns, and historical trend data are available for qualified prospects under NDA. Contact our team to request access.

Request benchmark data

Book a demo and we'll share detailed AMD benchmark results relevant to your language, region, and campaign type.