Solutions
Agatha Brand Perception Brand Equity Customer Segmentation Concept Testing Product Development Pricing Research Employee Experience All solutionsIndustries
Consumer Packaged Goods Food & Beverage Retail Healthcare Financial Services Technology All industriesPlatform
How it Works
See how Agatha turns survey data into decision-ready insights in minutes.
Features
Agatha AI
Coming soon · Your autonomous research analyst.
LiveSlides™
Live presentations that update as data comes in.
Solver Interface
The AI-powered respondent experience.
Libraries
Reusable question templates for faster study design.
QuantQual
Quant scale and qual depth — in one study.
MCP Server
Coming soon · Your studies, native in every AI tool.
AI Open-End™
Open-ended responses, quantified at scale.
Segments
Reusable audiences across the whole study.
Respondent Manager
Data-quality control for every study.
IdeaCloud™
Open-end verbatims, visualised by support strength.
HaloGraph™
Drill from themes to verbatims in one sunburst chart.
AugmentIQ™
Coming soon · Test your survey on AI respondents before launch.
Focus Groups
Live discussions and async bulletin boards.
Multilanguage
23 languages. One study. One dashboard.
Learn
Case Studies Real research, real results Blog Insights on market research FAQ Common questions answeredData Quality
Data Quality Reports Monthly evidence of research integrity Data Quality Our commitment to research integrity
Open-ended survey questions let people answer in their own words instead of picking from a list somebody else wrote. In a recent study of 795 consumers across the UK, Germany, and Spain, that single design choice decided what the research found. Run the same study with a fixed answer list and it would have reported that packaging design drives choice, and that lapsed users and people who never tried the category are one audience. Both conclusions would have been wrong.
This is the third installment of the Research Reality Check, a series about what happens to real research outcomes when the data going in is compromised. Story #1 looked at how cheap panels distort pricing research. Story #2 followed a global study through quality screening. This one covers a cleaner failure mode: a valid sample, answering in good faith, inside a question that never gave them room to say the true thing.
Short on time? Jump straight to the specific section:
Open-ended survey questions ask a respondent to supply the answer. Closed-ended questions ask them to select one. The gap between those two things is larger than it looks, because they measure different quantities.
A closed-ended question measures agreement with a hypothesis. Someone on the research team decided which attributes belong on the list, and respondents rate what they were handed.
An open-ended question measures salience. It captures what comes to mind first, in the language the customer uses, before any prompt shapes the answer.
For most of the history of survey research, teams paid a real penalty for that richness. Verbatims had to be read and coded by hand, which was slow, expensive, and prone to whoever built the code frame. So the default settled on closed-ended lists, and the industry learned to treat open-ends as color commentary next to the numbers.
That tradeoff no longer holds. When respondents evaluate each other’s answers during fielding, verbatim themes carry statistical support in the same way a scale item does. In this study, 61% of regular drinkers supported the theme that there was nothing they disliked about the product. That is a quantified finding about category satisfaction, sourced from language nobody on the research team wrote. GroupSolver’s AI Open-End approach handles this step, turning open-ended responses into measurable themes at scale.
The answer list is not a neutral container. It is a hypothesis, written before fielding, that the study then spends its budget confirming.
Three specific things go wrong.
Attributes you leave off the list never appear. If natural ingredients are not among the options, no amount of sample will surface them. The study can only rank what it was given, which means the ceiling on any finding is the imagination of whoever built the questionnaire.
Forced-choice grids manufacture importance. Rate ten attributes on a five-point scale and every one of them lands somewhere respectable. Nothing can score as a net negative, because the format has no way to express one. In this study, a max-diff style importance measure asked people to name both the most and least important attribute. Two attributes came back as net detractors, which a rating grid would have reported as mild positives.
Audience labels collapse groups that behave nothing alike. “Non-user” is a convenient bucket on a screener. Split it and the two halves point toward opposite budget decisions, as the results below show.
The study was run for a global beverage manufacturer trying to understand how consumers choose a bottled ready-to-drink beverage that competes with sodas and energy drinks, and why some of them walk away from it. One identical survey was fielded in three languages across the UK, Germany, and Spain. It collected 795 completed responses over a two-day field window in April 2026, at a 34% incidence rate and a median interview length of 14 minutes. The sample split into 621 regular drinkers, 65 infrequent or lapsed drinkers, and 109 people who had never tried the product, giving a full funnel from loyal users to the outright uninterested.
Taste carries the category. Packaging pulls against it.
When regular drinkers explained why they reach for this drink over a soda or an energy drink, the reasons clustered tightly:
Health positioning and convenience trailed well behind. The category’s edge over other cold beverages still runs through tasting good.
The importance model sharpened the point. Taste and flavor scored roughly 3.5 times higher than the next most important attribute, natural ingredients. Meanwhile “label and packaging design” and “glass or plastic” both came back as net detractors, meaning more people named them least important than most important. For a category that spends heavily on can and bottle creative, that is a humbling number. Packaging aesthetics are competing for the same brief space as the message that actually moves the category.
This is an out-of-home category before it is anything else.
Consumption skewed heavily toward moments away from the kitchen:
Several respondents were explicit that the product fills a gap when they are short on time or have no easy access to a fresh alternative. Anyone working on pack format or retail placement should design for the grab-and-go moment first and the stock-the-fridge moment second. That ordering falls out of the verbatims, not from a usage grid.
Lapsed drinkers and never-triers are two different problems.
People who had stopped drinking the product in the past year mostly cited price and a dislike of the cold format. Both are solvable with promotion and product work.
People who had never tried the category looked nothing like them. Only 7% said they were even somewhat likely to try it in the future, and their stated reasons came down to a flat lack of interest in the category itself. Winning lapsed drinkers back is a value and format problem with a plausible payback. Converting never-triers is a much longer road, and probably not worth chasing with the same budget line. Separating those groups cleanly is the kind of question customer segmentation exists to answer.
Pooled across the three countries, the results look like a single coherent European consumer. Cut by market, that consumer disappears.
A single pan-European positioning would have landed well in one market and undersold the category in at least one other. The averages were not wrong, exactly. They described a person who does not exist in any of the three countries.
Two design decisions made that comparison legitimate. The instrument was identical across all three markets, so nothing in the question wording explains the gaps. And it ran in three languages, so respondents wrote their open-ended answers in the language they actually think in. Fielding one study across multiple languages in a single dashboard is what makes market-level open-ended comparison possible at all.
The differences also showed up in what people volunteered, beyond how they rated the prompts they were shown. Sugar content came up unprompted in Spain. A closed-ended instrument built around a UK-first hypothesis would have collected a Spanish sample and reported that Spanish consumers care about flavor variety slightly less than the British do.
Open-ended survey questions have a reputation for being unquantifiable. That reputation comes from how they used to be processed, and it is fixable with sequencing and structure.
Ask the open question before you show the list. Order is not cosmetic. Once a respondent has seen ten attributes, their unprompted answer is contaminated by those ten attributes. Put the open-end first, then run the closed items to confirm.
Have respondents evaluate each other’s answers during fielding. This is the step that converts a pile of verbatims into theme support with statistical weight. Instead of coding after the fact, the sample itself tells you how widely each idea is held.
Field one instrument, translated, across every market. Rewriting a questionnaire per country produces country-level differences you cannot attribute to consumers. Translate the same study and keep the structure fixed.
Use an importance measure that can produce a negative. A max-diff style question forces a least-important pick, so an attribute can land below zero. This is how packaging design revealed itself as a detractor in a category that assumed it was a driver.
Split your non-users at the screener. Lapsed and never-tried belong in separate cells before analysis begins, because they justify different spending. Merging them produces an average recommendation that fits neither.
Keep the interview short enough that people still write. Open-ends collapse into one-word answers in a long grid-heavy survey. The median interview here ran 14 minutes, which is short enough that respondents were still composing real sentences at the end.
Calling “other, please specify” an open-ended question. It collects a trickle of responses from people already fatigued by the list, and those responses usually get coded back into the existing buckets anyway. The open-end has to be the primary question, with the list as the follow-up.
Coding verbatims into a pre-built code frame. If the frame was written before fielding, the study has reintroduced the fixed list one step later in the process. Let the themes emerge from what respondents actually wrote, then measure support for them.
Reading the pooled multi-market average. A regional average can look healthy and coherent while describing nobody. Always cut by market before you write the positioning line.
Treating “nothing I dislike” as a non-answer. It reads like a respondent being lazy. In this study 61% of regular drinkers supported that theme, which is a substantive finding about how settled the category feels to its core users. A closed-ended dislike battery would have forced those people to invent a complaint.
Assuming a larger sample fixes it. Sample size improves precision on the options you offered. It does not add the option you left out.
Open-ended survey questions measure what is salient to a customer. Closed-ended questions measure agreement with the hypothesis the research team wrote before fielding.
Across 795 consumers in the UK, Germany, and Spain, taste and flavor scored roughly 3.5 times higher in importance than the next attribute, while packaging design and container material came back as net detractors.
The category is consumed away from home: 60% while traveling and 54% at work, which points pack format and retail placement toward grab-and-go.
Lapsed drinkers cited price and cold format, both solvable. Only 7% of never-triers said they were even somewhat likely to try the category, making them a much more expensive audience to pursue.
Market-level cuts split three ways: flavor variety in the UK, natural ingredients and sugar in Spain, price and convenience in Germany. A single European positioning would have undersold at least one of them.
Verbatim themes can be quantified when respondents evaluate each other’s answers during fielding. In this study, 61% of regular drinkers supported the theme that there was nothing they disliked.
The reality check in this story is less dramatic than the previous two. No fraud, no mass removal. Just 795 real people answering carefully inside a design that would have thrown away the most useful thing they had to say. Open-ended survey questions are what turned a routine category read into a brief that changed how a manufacturer thinks about packaging spend and which non-users are worth chasing.
This is Study #3 of the Research Reality Check. If it gave you something to think about, more are on the way.
Stay curious. Ask why.
PEOPLE ALSO READ

GroupSolver insights

GroupSolver insights

GroupSolver insights
Book a 30-minute demo with our research team. No deck, no pitch — just the platform answering your questions.
Book a 30-min demoNo credit card · we reply in 24h