Solutions
Agatha Brand Perception Brand Equity Customer Segmentation Concept Testing Product Development Pricing Research Employee Experience All solutionsIndustries
Consumer Packaged Goods Food & Beverage Retail Healthcare Financial Services Technology All industriesPlatform
How it Works
See how Agatha turns survey data into decision-ready insights in minutes.
Features
Agatha AI
Coming soon · Your autonomous research analyst.
LiveSlides™
Live presentations that update as data comes in.
Solver Interface
The AI-powered respondent experience.
Libraries
Reusable question templates for faster study design.
QuantQual
Quant scale and qual depth — in one study.
MCP Server
Coming soon · Your studies, native in every AI tool.
AI Open-End™
Open-ended responses, quantified at scale.
Segments
Reusable audiences across the whole study.
Respondent Manager
Data-quality control for every study.
IdeaCloud™
Open-end verbatims, visualised by support strength.
HaloGraph™
Drill from themes to verbatims in one sunburst chart.
AugmentIQ™
Coming soon · Test your survey on AI respondents before launch.
Focus Groups
Live discussions and async bulletin boards.
Multilanguage
23 languages. One study. One dashboard.
Learn
Case Studies Real research, real results Blog Insights on market research FAQ Common questions answeredData Quality
Data Quality Reports Monthly evidence of research integrity Data Quality Our commitment to research integrity
Survey data quality problems rarely announce themselves. The file arrives on schedule, the crosstabs balance, the charts look decisive, and none of that tells you whether real people answered honestly. A survey data quality checklist replaces the eyeball test with twelve specific questions you can ask about the dataset already sitting on your desk, and about the supplier who sold it to you.
Short on time? Jump straight to the specific section:
A survey data quality checklist is a fixed set of questions applied to a dataset and its sample source to determine whether the results are reliable enough to base a decision on.
It is not a satisfaction score for your vendor. It tests four things:
Provenance: where every completed interview came from, and who else touched it
Authenticity: whether the people behind the responses are real, unique, and actually who they claimed to be
Effort: whether respondents engaged with your questions or pattern-matched their way to the incentive
Integrity of the conclusion: whether your headline finding survives the removal of everything you flagged
Most published guidance on how to evaluate survey data quality is written for the procurement moment: you’re picking a panel, here’s what to ask. That’s useful once a year. The checklist below is built so that at least ten of the twelve questions can be answered about data you already hold, without calling anyone.
Here’s the uncomfortable part. The two checks nearly every research team runs (speeder flags and attention traps) are the two that perform worst.
When Pew Research Center compared six online sample sources across more than 60,000 interviews, it found that crowdsourced and opt-in panel respondents were considerably more likely than address-recruited respondents to submit bogus data, and that two standard quality checks — one for speeding, one for attention — failed to catch most of the respondents ultimately flagged as bogus. The screens were running. They just weren’t finding anything.
The failure mode is worse than random noise. Pew found that bogus respondents don’t answer at random, they lean toward positive answer choices, which biases estimates upward rather than simply adding scatter. Random error widens your confidence interval. Systematic acquiescence moves your number in a direction, and moves it quietly.
The scale of what gets caught when someone looks properly is easy to underestimate. In GroupSolver’s published monthly data quality reports, the share of completed responses removed as bots, fraud, or low-quality answers has ranged from 41% to 66% across the last six reporting months, with a three-month rolling average of 62% as of July 2026. In July alone, two studies run through the identical process came back at 43% and 63%. Same standard, same month, twenty points apart.
That spread is the real lesson. A removal rate isn’t a verdict on a research vendor’s rigor, it’s a reading on that particular sample. Which is exactly why you need to know your own number before you can interpret anyone’s findings, including your own.
Panel supplier quality is the input variable most teams treat as fixed. It isn’t.
Where did every complete actually come from, and can your supplier name the sources? Many samples are assembled through aggregators that blend traffic from dozens of underlying panels, routers, and exchanges. If your supplier can only tell you the name of the marketplace, you don’t know your sample source; you know your invoice. Ask for the composition, in writing, per study.
What percentage of your completes were removed, and why? This is the single most diagnostic question in the checklist, and you can answer it about data you already own. If the answer is “we didn’t remove any,” that is not a clean sample, that’s an unexamined one. If the answer is “about 12%,” ask what the categories were.
Do removal rates differ by supplier, and is anyone tracking that over time? One study tells you very little. Six months of removal rates by source tells you which supplier is quietly costing you money. Most teams never build this view because no one owns it, and the real cost of cheap survey panels shows up in pricing and concept work long before anyone traces it back to sample.
Survey fraud detection is less about catching sophisticated attackers and more about noticing people who claimed things that can’t be true.
How many respondents claimed something almost nobody is? Pew ran this test directly: in a February 2022 experiment, 12% of opt-in respondents under 30 said they were licensed to operate a nuclear submarine — a qualification held by effectively 0% of Americans. Plant one low-incidence claim in your screener. The share who select it is your fraud floor for that sample.
Would you catch an AI-written open end if you saw one? Generated text is fluent, on-topic, and structurally tidy, which is precisely what makes it slip past reviewers trained to look for gibberish. The signal is uniformity across respondents, not incoherence within one. This is a fast-moving detection problem, and AI-generated survey responses now require different checks than the ones written for 2019-era bad survey data.
What happens when the same person qualifies twice? Duplicates enter through device sharing, VPN cycling, and panelists holding accounts across multiple suppliers. Ask what your deduplication actually keys on, panelist ID alone is the weakest version of this check.
Respondent quality checks fail most often when they measure the wrong thing: how a person answered, rather than whether the answer addressed the question.
Are your open ends answering the question, or just producing text? Pew’s study found bogus responses included answers with no bearing on the question asked, along with uniformly positive answers regardless of the topic. A response can be well-formed, correctly spelled, the right length — and still be about nothing. Read fifty at random before you trust a single coded theme.
Does anyone straight-line, and did you look before or after weighting? Weighting is applied to correct demographic imbalance. Applied to a sample containing flat-lined grids, it can amplify a small cluster of low-quality survey responses into a visible pattern. Check the raw file.
How fast is “too fast,” and who set that threshold? Speeder thresholds are usually inherited, rarely revisited, and catch far less than teams assume. Set the threshold against your own median completion time for this instrument, then treat it as one weak signal among several rather than a pass/fail gate.
These three are where data quality in market research stops being a methods conversation and becomes a decision-making one.
What does your headline finding look like with the flagged cases removed? Run the number both ways. In one of Pew’s opt-in panels, presidential approval fell two points (from 42% to 40%) once bogus cases were excluded. Two points is trivial in a tracker and decisive in a concept test with a go/no-go threshold at 40%.
Does your base size still support the cut you’re presenting? Removals aren’t distributed evenly. If 30% of your sample came out and most of it sat in one age bracket, the segment slide you were planning to lead with may now rest on n=60. Recheck every subgroup base after cleaning, not before.
Would you be comfortable publishing your removal rate next to your headline number? If the answer is no, you already know something about the data that your audience doesn’t. That instinct is worth taking seriously.
You don’t need a new study to start. Run this against the last one you fielded.
Pull the raw, unweighted file. Not the cleaned deliverable, not the dashboard export. Cleaning decisions someone else made are part of what you’re auditing.
Answer questions 2, 4, 7, and 8 first. These four are the highest-yield and require nothing from your supplier. Together they’ll tell you within an hour whether the rest of the checklist is urgent or routine.
Rerun your top three findings with flagged cases excluded. Document the delta on each. If nothing moves, you have a genuinely robust result, and now you can say so with evidence.
Take the gaps to your supplier as questions, not accusations. Questions 1, 3, and 6 are the supplier-facing ones. Framed as a standing quarterly review rather than a complaint, they change behavior far more reliably.
The teams that get furthest with this stop treating it as a periodic audit and build it into the workflow. That’s the design principle behind Agatha — detection running before data lands in analysis, followed by a manual pass for coherence and relevance, so the checklist isn’t a rescue operation performed after the deck is built.
None of these twelve questions require new tooling, a bigger budget, or a change in supplier. They require someone to look at data everyone else assumed was fine. That’s the whole gap between decision-grade data and confident-looking noise, not sophistication, just the willingness to check.
Run the survey data quality checklist against your most recent study this week. Whatever you find, you’ll know something your competitors’ research teams don’t know about theirs. Twelve questions, one for each of the twelve years we’ve spent asking them.
A survey data quality checklist tests provenance, authenticity, effort, and whether your conclusion survives cleaning.
Speeder and attention checks miss most bad respondents — running them is not the same as catching anything.
Bogus respondents skew positive, so they bias your estimate in a direction rather than adding random noise.
Removal rates vary widely between suppliers and between studies in the same month; only tracking over time makes the number meaningful.
Ten of these twelve questions can be answered about data you already hold, today, without contacting a vendor.
The decisive test is whether your headline finding changes once flagged cases are removed.
The lesson from this checklist isn’t that most research is broken, it’s that survey data quality is the hidden variable determining whether a study informs a decision or quietly misleads it. None of these twelve questions require sophistication, just the willingness to check before you present.
And if you wonder why there are twelve questions, it’s because this year marks twelve years of applying AI to survey research, one question for each year we’ve spent asking why the data looks the way it does.
Stay curious. Ask why.
Book a 30-minute demo with our research team. No deck, no pitch — just the platform answering your questions.
Book a 30-min demoNo credit card · we reply in 24h