Strong separation in a controlled benchmark study.
SurveyShell was tested through 612 deliberately varied survey journeys whose operating method was known from the study setup.
Method summary
- Fieldwork dates: 20 July to 4 August 2026
- Method: online test survey of 19 pages, completed on desktop, mobile and tablet devices
- Sample: 612 controlled test journeys — 203 human-operated, 208 Indralo-built automation, 201 AI-agent-operated; deliberately varied, not a representative sample
- Weighting: none
- Scoring engine: v2-engine-0.7.0-bootstrap
Human-operated and AI-agent-operated journeys
Almost every human-operated journey was classified Likely human: 202 of 203.
Every AI-agent-operated journey was classified Unlikely human: 201 of 201.
| Controlled test journey | Likely human | Needs review | Unlikely human | Total |
|---|---|---|---|---|
| Human-operated | 202 | 1 | 0 | 203 |
| AI-agent-operated | 0 | 0 | 201 | 201 |
| Total | 202 | 1 | 201 | 404 |
The unit is a controlled survey journey, not a unique person or AI system. The same human operators and AI tools could complete more than one journey.
Indralo-built test automation
| Controlled test journey | Likely human | Needs review | Unlikely human | Total |
|---|---|---|---|---|
| Indralo-built test automation | 9 | 7 | 192 | 208 |
These journeys were generated using automation created and operated by Indralo. They provide an engineering stress test, not independent validation against bots observed in commercial fieldwork. The same scripts could complete more than one journey.
The difficult cases stay on the page
Publishing all three outcomes shows both the strong result and the automation that remained difficult to distinguish. Because Indralo created and operated this automation, it is an engineering stress test rather than independent fieldwork validation.
Deliberately varied controlled journeys
Unique survey links were allocated to three groups: human operators, Indralo-built test automation and browser-capable AI agents. Each link was tracked, so every journey's type came from how its link was used, not from SurveyShell's result.
Human operators repeated the survey across desktop, mobile and tablet devices using mouse, touchscreen, pen and keyboard input. The recorded set included macOS, Windows, iOS and Android devices, along with deliberately varied survey-taking behaviours.
Several browser-capable AI agents completed repeated journeys using varied instructions and personas. Indralo-built automation ranged from simple scripts to stealthier, vision-based and more human-like test automation.
How to read the results
Human-operated journeys
Of 203 human-operated test journeys, 202 were classified Likely human, one Needs review and none Unlikely human.
AI-agent-operated journeys
All 201 AI-agent-operated journeys in this benchmark study were classified Unlikely human. This is the observed result for this test group, not a claim that every current or future AI agent will receive the same result.
Indralo-built test-automation journeys
Of 208 in-house automation journeys, 192 were classified Unlikely human, seven Needs review and nine Likely human. Publishing all three outcomes shows both the strong result and the automation that remained difficult to distinguish. Because Indralo created and operated this automation, it is an engineering stress test rather than independent fieldwork validation.
What this supports
A promising controlled result
Under the tested conditions, the human-operated and AI-agent-operated groups separated clearly: no human-operated journey reached Unlikely human, while every AI-agent-operated journey did.
The separate in-house automation table includes the difficult cases in full: nine journeys were classified Likely human and seven Needs review.
What this does not prove
Not a real-world accuracy rate
This was an intentionally constructed and repeated controlled benchmark study, not a random or representative sample of commercial fieldwork. It was not independently conducted or audited, and it does not establish a general accuracy, false-positive or population fraud-detection rate.
The test cannot represent every survey, device, accessibility method, respondent population, network condition or future automation strategy. It also does not validate SurveyShell's separate Quality feature, which requires different evidence.
Human Likelihood remains advisory. The bands describe the recorded interaction, not the person: automation-consistent interaction can also come from a human using assistive tools such as voice control.
Varied live fieldwork
The next step is to evaluate SurveyShell across varied live studies, devices, sample sources and survey platforms, with results reviewed alongside each research team's existing checks.