Benchmarks
AMD performance. Measured honestly.
Our benchmark methodology, the metrics we track, the languages we test, and the limitations we acknowledge. Full data available under NDA.
What We Benchmark
AMD accuracy across languages
DialerBee's AI AMD uses transcript classification on the first 3 seconds of call audio to determine whether a human or an answering machine has picked up. We benchmark this system across multiple languages, carrier environments, and campaign types to measure real-world accuracy.
Unlike vendors who report lab-condition accuracy, our benchmarks reflect production data with all the noise, carrier variability, and regional differences that come with real outbound campaigns.
Methodology
How we measure accuracy
Transcript Classification
The AMD model classifies a real-time transcript of the first 3 seconds of audio. It reads words and phrases, not audio waveforms or beep patterns.
Agent Override as Ground Truth
When agents disagree with the AMD classification, they flag it with one click. These overrides serve as ground truth labels for measuring accuracy.
Continuous Evaluation
Benchmarks are not one-time tests. We continuously evaluate accuracy against the stream of agent overrides across all production campaigns.
Metrics Tracked
Five metrics that matter
Accuracy
Overall correct classification rate across all calls.
False Positive Rate
Human calls incorrectly classified as machines.
False Negative Rate
Machine calls incorrectly classified as human.
Detection Latency
Time from call pickup to AMD classification decision.
Agent Override Rate
Percentage of calls where agents correct the AMD decision.
Benchmark Summary
Public summary from internal pilots
Detailed per-language, per-carrier, and per-region data available under NDA for qualified prospects.
| Metric | Public Summary | Notes |
|---|---|---|
| Calls evaluated | 10,000–50,000 range | Across internal pilot campaigns |
| Date range | Q1–Q2 2026 | Updated as new pilot data becomes available |
| Languages tested | 9 | EN, AR, ES, FR, IT, DE, TR, HI, UR |
| Regions tested | MENA, Europe, North America | Carrier mix varies by region |
| Campaign types | Collections, BPO, sales, insurance | Results segmented where possible |
| AMD false positive rate | Below 3% in selected pilots | Human incorrectly classified as machine |
| AMD false negative rate | Below 5% in selected pilots | Machine incorrectly classified as human |
| Median detection latency | Sub-1 second | Pickup to AMD decision |
| P90 detection latency | Under 1.5 seconds | Pickup to AMD decision |
| Agent override rate | 2–5% range | Manual correction rate in pilot campaigns |
| Contact-rate lift | Up to 4x vs manual dialing | In pilot workflows; results vary |
| Idle-time reduction | 30–45% range | In selected pilot campaigns |
Based on internal pilot benchmarks. Results vary by campaign type, list quality, carrier, region, language, pacing configuration, and agent workflow. Full per-language and per-carrier breakdowns available under NDA.
Languages Tested
Nine languages, real pilot data
EN
English
AR
Arabic
ES
Spanish
FR
French
IT
Italian
DE
German
TR
Turkish
HI
Hindi
UR
Urdu
Arabic benchmarks include Gulf, Levantine, and Egyptian dialect variants. Hindi and Urdu benchmarks include regional variant coverage. Each language is benchmarked independently with carrier-specific and region-specific breakdowns.
Limitations
What benchmarks cannot tell you
Results vary by carrier. Different carriers produce different voicemail greetings, audio quality, and connection timing.
Results vary by campaign type. Collections campaigns have different pickup patterns than sales or survey campaigns.
Results vary by region. Voicemail greeting styles, carrier infrastructure, and network conditions differ across markets.
Results vary by list quality. Stale numbers, wrong numbers, and disconnected lines affect AMD performance differently.
Benchmarks reflect aggregate performance. Individual campaign results may be higher or lower than reported averages.
Full benchmark data available under NDA
Detailed benchmark reports with per-language accuracy tables, carrier-specific breakdowns, and historical trend data are available for qualified prospects under NDA. Contact our team to request access.
Request benchmark data
Book a demo and we'll share detailed AMD benchmark results relevant to your language, region, and campaign type.