In November 2025 Sean Westwood of Dartmouth published a paper in PNAS describing an autonomous AI survey respondent built from a prompt of about 500 words. Across 43,000 tests it passed 99.8% of attention checks. By the university's description, it could complete a survey for around five cents, against the $1.50 or so a human respondent typically earns.
Most of the commentary treated this as the end of online research. We read it differently. The paper shows that one category of defence, questions inside the questionnaire, no longer tells you whether a respondent is a person. It says much less about the other places a session can be checked.
Disclosure before going further: ResearchDart sells pre-survey screening. Weigh the argument accordingly.
Why attention checks stopped working
An attention check asks the respondent to prove they read the question. "Select 'somewhat disagree' for this row." For a decade that was a reasonable filter for bored humans and cheap scripts, because both failed it often.
A language model reads every question. It never gets bored. It follows the instruction in the row. Asking it to prove it is paying attention is asking it to do the one thing it does perfectly. Logic puzzles, reverse-coded items and consistency pairs fail for the same reason: they test comprehension, and comprehension is what these models are good at.
Keep your attention checks. They still catch inattentive people, and inattentive people remain the bigger share of bad data in most studies. Just stop treating a pass as evidence of a human.
What an agent cannot easily hide
An agent answering a survey is software running somewhere, operated by someone, usually at volume. Those three facts leave traces outside the questionnaire.
It runs in an automated browser. Agents drive a browser through automation interfaces. Some leave obvious marks, such as headless user agents. Better ones need deeper device and behaviour signals to spot. CloudResearch published results in June 2026 showing that agent-driven cursors move in straight lines and jump to targets, where human cursors overshoot and correct. That is a behavioural signal, and it lives in the browser, not in the answers.
It runs on infrastructure. Volume fraud tends to come from hosting providers, VPNs and proxy networks, because that is where you can run hundreds of sessions cheaply. Genuine panel members mostly connect from home broadband and mobile networks.
It is operated by someone with many accounts. Five cents a complete only pays if you do it thousands of times, which means many panel memberships and many supplier IDs. Multi-accounting leaves device and identity patterns that repeat across sessions.
None of these is perfect on its own. Together they are much harder to fake than a fluent open end.
Where to check, in order
- Before the survey loads. Device, network, automation and duplicate checks at the entry link. A respondent stopped here costs the buyer nothing and never contaminates the data. This is where the Insights Association and GDQ benchmark saw removals shifting in its 2025 wave.
- While the survey runs. Behavioural telemetry inside the questionnaire, where your survey platform or a specialist tool supports it. This catches agents that pass the entry checks, at the cost of a heavier integration.
- After the survey. Open-end review, timing analysis, straight-lining. This is still necessary, but it is the most expensive place to find fraud, because you have already paid for the complete and have to reverse it.
If you only add one layer this year, add the first one. Our reasoning for that is in what pre-survey screening catches and what it misses.
What not to rely on
AI text detectors on open ends. They produce confident scores with poor accuracy, and they flag non-native English writers disproportionately. The GDQ code frame does include a reason for AI-completed open ends (code 15). Use it when the evidence is plain, such as identical phrasing across respondents or answers that quote the question back in a chatbot's register, not because a detector said 80%.
Longer screeners. Adding more qualifying questions does not stop an agent that reads well. It just raises LOI for real respondents.
Captcha on the entry page. Modern agents solve most of them, and every extra step costs genuine completes on mobile.
Questions to ask your sample suppliers
These are fair questions in 2026. A supplier that cannot answer them is telling you something.
- Do you screen sessions before they reach my survey, and with what: your own checks, a third-party tool, or both?
- What share of your traffic did that screening remove last quarter, for studies like mine?
- Do you block datacenter, VPN and proxy traffic, and what do you do about corporate VPNs on B2B studies?
- Do you detect one person holding several accounts across your panel or your sub-sources?
- When I reverse completes, do you accept reasons coded with the GDQ removal codes, and what happens to the sub-source afterwards?
- Are your entry links and redirects signed or server-to-server?
What we changed because of this
We moved screening in front of the questionnaire and made it a per-project setting with a monitor mode, so a team can measure its own fraud rate before blocking anyone. We map every router decision to a GDQ code so pre-survey blocks and post-survey reversals land in one report. And we are explicit that the built-in automation check only catches careless automation. Determined agents need a proper screening provider. Details are on the pre-survey screening page.
Online sample is not over. But a study that relies only on questions inside the survey to prove its respondents are human is relying on the one defence that stopped working.