Benchmarks

AMD performance. Measured honestly.

Our benchmark methodology, the metrics we track, the languages we test, and the limitations we acknowledge. Full data available under NDA.

What We Benchmark

AMD accuracy across languages

DialerBee's AI AMD uses transcript classification on the first 3 seconds of call audio to determine whether a human or an answering machine has picked up. We benchmark this system across multiple languages, carrier environments, and campaign types to measure real-world accuracy.

Unlike vendors who report lab-condition accuracy, our benchmarks reflect production data with all the noise, carrier variability, and regional differences that come with real outbound campaigns.

Methodology

How we measure accuracy

Transcript Classification

The AMD model classifies a real-time transcript of the first 3 seconds of audio. It reads words and phrases, not audio waveforms or beep patterns.

Agent Override as Ground Truth

When agents disagree with the AMD classification, they flag it with one click. These overrides serve as ground truth labels for measuring accuracy.

Continuous Evaluation

Benchmarks are not one-time tests. We continuously evaluate accuracy against the stream of agent overrides across all production campaigns.

Metrics Tracked

Five metrics that matter

Accuracy

Overall correct classification rate across all calls.

False Positive Rate

Human calls incorrectly classified as machines.

False Negative Rate

Machine calls incorrectly classified as human.

Detection Latency

Time from call pickup to AMD classification decision.

Agent Override Rate

Percentage of calls where agents correct the AMD decision.

Benchmark Summary

Public summary from internal pilots

Detailed per-language, per-carrier, and per-region data available under NDA for qualified prospects.

Metric Public Summary Notes
Calls evaluated 10,000–50,000 range Across internal pilot campaigns
Date range Q1–Q2 2026 Updated as new pilot data becomes available
Languages tested 9 EN, AR, ES, FR, IT, DE, TR, HI, UR
Regions tested MENA, Europe, North America Carrier mix varies by region
Campaign types Collections, BPO, sales, insurance Results segmented where possible
AMD false positive rate Below 3% in selected pilots Human incorrectly classified as machine
AMD false negative rate Below 5% in selected pilots Machine incorrectly classified as human
Median detection latency Sub-1 second Pickup to AMD decision
P90 detection latency Under 1.5 seconds Pickup to AMD decision
Agent override rate 2–5% range Manual correction rate in pilot campaigns
Contact-rate lift Up to 4x vs manual dialing In pilot workflows; results vary
Idle-time reduction 30–45% range In selected pilot campaigns

Based on internal pilot benchmarks. Results vary by campaign type, list quality, carrier, region, language, pacing configuration, and agent workflow. Full per-language and per-carrier breakdowns available under NDA.

Languages Tested

Nine languages, real pilot data

EN

English

AR

Arabic

ES

Spanish

FR

French

IT

Italian

DE

German

TR

Turkish

HI

Hindi

UR

Urdu

Arabic benchmarks include Gulf, Levantine, and Egyptian dialect variants. Hindi and Urdu benchmarks include regional variant coverage. Each language is benchmarked independently with carrier-specific and region-specific breakdowns.

Limitations

What benchmarks cannot tell you

Results vary by carrier. Different carriers produce different voicemail greetings, audio quality, and connection timing.

Results vary by campaign type. Collections campaigns have different pickup patterns than sales or survey campaigns.

Results vary by region. Voicemail greeting styles, carrier infrastructure, and network conditions differ across markets.

Results vary by list quality. Stale numbers, wrong numbers, and disconnected lines affect AMD performance differently.

Benchmarks reflect aggregate performance. Individual campaign results may be higher or lower than reported averages.

Full benchmark data available under NDA

Detailed benchmark reports with per-language accuracy tables, carrier-specific breakdowns, and historical trend data are available for qualified prospects under NDA. Contact our team to request access.

Request benchmark data

Book a demo and we'll share detailed AMD benchmark results relevant to your language, region, and campaign type.