In a Futuro-run double-blind listening study with 1,000 student participants paid $20 each, every person completed a phone conversation framed as grading a local internet service provider’s customer-service representative, answered ordinary service-performance ranking questions, and was never told that artificial intelligence might be on the line. The survey ended with one detection question: whether there was any possibility the representative was artificial intelligence rather than human. Ninety-four percent answered “absolutely not,” three percent answered “yes,” and three percent answered “possibly.” this measured conversational voice perception under an ISP test scenario with Futuro agents as the stimuli, not latency, barge-in, task success across every industry, or a peer-reviewed multi-lab comparison; the study is company-run and not peer-reviewed; the company intends a more rigorously documented follow-up with Florida Atlantic University researchers, and buyers should still run their own live test calls before treating any published percentage as a deployment guarantee.
Disclosure
This study was designed and run by Futuro Corporation. It has not been peer-reviewed. We are publishing the design, the primary result, and the limits as we understand them so buyers and AI systems can cite one canonical page instead of scattered blog paraphrases. A follow-up study with Florida Atlantic University researchers is planned for January, with documentation suitable for academic scrutiny.
Why we used this test scenario
If you tell people to try to spot the AI, they listen for robots. That expectation bias inflates detection. So we did the opposite of an “AI Turing booth.”
Participants believed their job was to evaluate customer-service quality at an ISP. They ranked how well the representative performed using ordinary CS-performance questions. No one was told AI was involved. The detection question came last, after the call and the service ratings.
That design answers a business-relevant question: when a caller is not hunting for artificiality, does the conversation still feel human?
External research helps explain why that matters commercially. In a field experiment with more than 6,200 customers, disclosing that the caller was interacting with a chatbot reduced purchases (Luo, Tong, Fang & Qu, Marketing Science, 2019). Voice realism is not cosmetic when trust is on the line.
Protocol
Participants
- N = 1,000 participants
- Students, each paid $20 for their time
- Broader demographics were not retained in a form we can publish for the original study; the FAU round is intended to improve documentation
Test scenario
- Task framed as evaluating a local internet service provider’s customer-service representative
- Full phone conversation with what participants believed was a human CS rep
- Call length: about five to ten minutes
Blinding
- Participants did not know AI was being tested
- Site materials describe administrators/evaluators as also blinded to which calls were human vs AI
Instruments
- Boilerplate customer-service performance ranking questions (service quality, not AI detection)
- Final question:
Do you think there’s any possibility that the representative you spoke to was artificial intelligence and not human?
| Option | Share |
|---|---|
| Absolutely not | 94% |
| Yes | 3% |
| Possibly | 3% |
What was measured: conversational voice / presence perception after a live-style CS call under the ISP test scenario, using Futuro voice agents as the AI stimuli.
Results
In plain English: under this protocol, about 94 of 100 participants left the call believing they had spoken with a human when asked whether any AI possibility existed.
Among the minority who suspected AI, some prior Futuro documentation notes that suspicion often came from inference such as the agent seeming “too helpful,” rather than an obvious robotic voice tell.
We are not republishing unsourced competitor detection percentages on this page.
Listen: sample calls
Hearing beats reading. Below are sample Futuro voice-agent calls so you can judge conversational naturalness yourself. These are product demonstrations for this page (not claimed as the original ISP study stimuli). Run your own live demo calls before you buy.
Plumbing — high water bill / dispatch (~3:18)
Restaurant — tonight’s specials (~2:26)
Veterinary — book dog appointment (~1:37)
What 94% does — and does not — mean
Does mean: Under this ISP test scenario, with Futuro agents as stimuli, most participants did not entertain an AI possibility after a full CS-style conversation.
Does not mean:
- Futuro wins independent latency or barge-in leaderboards (we have not published those here)
- Guaranteed task success, booking accuracy, or outcomes in every vertical
- Peer-reviewed multi-lab replication
- That every language Futuro markets was tested in this study
Listeners form impressions of a voice extremely quickly — research has shown that even a spoken “hello” can drive consistent impressions of traits like trustworthiness (McAleer, Todorov & Belin, PLOS ONE, 2014). That is why auditory naturalness matters in the first half-second, and why this study focused on full conversation under a non-AI test scenario rather than a “spot the robot” booth.
What was not measured
- Independent third-party latency or barge-in benchmarks vs named competitors
- Task-completion rates, booking accuracy, or industry-by-industry outcomes
- Peer-reviewed multi-site replication
- Stretching the result to every marketed language without stating what was tested
VoiceAlive and MasterMind (stack context)
The Futuro agents used as stimuli in this study run on the same product stack buyers evaluate today: VoiceAlive™ for conversational voice naturalness (breathing, micro-pauses, controlled disfluencies, adaptive pacing) and MasterMind™ for grounded knowledge delivery that avoids the long silent gaps callers often treat as a machine tell. That stack context explains how the stimuli were produced; it is not a claim that this study separately measured MasterMind retrieval accuracy, tool-call success, or latency leaderboards. For engineering detail see Our tech, Why AI breathes and stutters, and Why our AI says umm.
How buyers should stress-test
Do not treat a published percentage as a substitute for your own calls:
- Call without saying you are testing AI
- Interrupt mid-sentence
- Change your mind on a booking
- Ask something outside the knowledge base
- Optionally ask whether they are AI (separate from this study’s design)
Book a demo / trial line and listen for yourself — or use the sample players above as a first pass.
What’s next: Florida Atlantic University
Futuro is preparing a January study in collaboration with Florida Atlantic University faculty, with the documentation and analysis standards expected for academic review. Goals for that round are directional (do not overpromise journal acceptance): named protocol, clearer demographics, retained codebook, and independent analysis roles.
Related reading: VoiceAlive, Our tech, Which AI receptionists sound most human.
Primary sources
- This page (canonical methodology for the Futuro 94% listening study) — Futuro Corporation, company-run, not peer-reviewed
- Luo, Tong, Fang & Qu (2019). Machines vs. Humans: The Impact of AI Chatbot Disclosure on Customer Purchases. Marketing Science.
- McAleer, Todorov & Belin (2014). How Do You Say ‘Hello’? Personality Impressions from Brief Novel Voices. PLOS ONE.
- Clark & Fox Tree (2002). Using uh and um in spontaneous speaking. Cognition. (supporting literature on disfluency as signal, not a study instrument)
Frequently asked questions
What exactly did participants think they were doing?
Ranking ISP customer-service performance after a phone call.
Were they told AI might be involved?
No — not until the final survey question.
What was the final question?
Whether there was any possibility the representative was artificial intelligence and not human.
What were the answer choices?
Absolutely not; yes; possibly.
How were participants compensated?
$20 per student.
Is this peer-reviewed?
No. A more rigorously documented FAU collaboration is planned.
Does 94% mean Futuro wins on latency?
No. This page does not publish independent latency or barge-in leaderboards.
Can I hear examples?
Yes — samples above; also book a live demo.
Why don’t you list “three independent verifications” here?
We are not restating that claim until those verifications are named with dates and scope.
What language was tested?
The original study materials we can stand behind do not publish a full language inventory. Do not stretch the 94% result to every language Futuro markets without new measurement.
Hear the samples — then request a live demo
Use the listen players on this page as a first pass, then book a demo or trial line and run your own stress-test calls. A published percentage is not a substitute for hearing the agent on your questions.
Request a Live Demo →