Answer options that hold up: designing multiple-choice questions
The list of answers under a question decides what people can tell you. Here is how to build answer options that are exclusive, exhaustive and honest, and why select-all-that-apply quietly undercounts.
By SurveyLane · The team building SurveyLane
Most advice about survey questions is about the question. The answer list gets written in thirty seconds afterwards, and that is where a lot of the damage happens. The options you offer decide which answers are possible and which ones never get counted. Get the list wrong and a well-worded question still produces numbers you cannot defend.
The answer list is part of the question
Respondents read the question, look at the boxes and pick the one that fits best. If a reasonable answer is missing, they take the nearest box. If two overlap, they pick one at random. If there are fourteen, they read the first few and stop.
None of this shows in your export, which is why the post on writing better survey questions treats the options as part of the item.
Pew Research Center puts the core rule in one sentence: closed-ended questions should include all reasonable responses, so the list is exhaustive, and the categories should not overlap, so they are mutually exclusive.
Mutually exclusive: no answer should fit two boxes
Overlap is the most common mistake. Numeric ranges are the classic case.
How old are you?
18-25
25-35
35-45
45 or older
A 25-year-old fits two boxes. So does a 35-year-old. Some pick the lower band, some the higher, and your bands blur at exactly the edges you care about. The fix is boring: 18-24, 25-34, 35-44, 45 or older.
Categories overlap too. "How did you hear about us?" with "Social media", "Instagram", "A friend" and "Online" gives someone who saw a friend's Instagram post four defensible answers. Build the list along one dimension, say the channel, and ask a second question for a second dimension.
Frequencies overlap more quietly. "Daily" and "Several times a week" blur for someone who uses the product five days out of seven. Anchor them to counts, like "5 or more days a week".
Exhaustive: every respondent needs a box
Everyone must find an option that is true for them. Otherwise they pick the closest one, and it absorbs answers that don't belong there.
Pew has a clear example of how hard a list steers. After the 2008 US presidential election, one group was asked which issue mattered most to their vote and was read five options. Another group got the same question with no list. When the economy was offered, 58% chose it. In the open version, only 35% volunteered it. In the closed version, 8% gave an answer outside the five offered. In the open version, 43% named something that wasn't on the list at all.
So the list shaped the opinions it was supposed to measure. Leaving an option out makes it close to invisible, because few people volunteer what you didn't suggest. If your list of cancellation reasons lacks "my company changed tools", you will never learn how many people left for that reason.
"Other, please specify" catches less than you hope
The usual patch for an incomplete list is an "Other" box with a text field. Keep it, but the 2008 numbers show its limit: people who get a list mostly choose from it, and "Other" only collects answers from those motivated enough to type.
Use it as an early warning. Read what comes in during the first days of fielding. If the same answer keeps turning up, it was a missing option, and the count you see is an undercount. Add it for the next wave. Don't recode it and present it as if everyone had been offered it.
Put "Other" at the bottom, after the real options. And make its text field optional, because a mandatory "please specify" turns a quick tick into a writing task and pushes some people to tick something else.
Why "select all that apply" quietly undercounts
"Which of the following have you experienced? Select all that apply." The checklist is compact and feels efficient. It also underreports.
Pew tested it. On its online American Trends Panel, from July 30 to August 12, 2018, among 4,581 US adults, half the sample saw a list of events as a select-all checklist. The other half answered each event as a separate yes/no question. In the report by Arnold Lau and Courtney Kennedy, published May 9, 2019, the estimates were on average 8 percentage points higher with forced choice.
Some items moved a lot more. Being denied health insurance went from 13% in the checklist to 19% with forced choice. Losing a job or struggling to find another went from 51% to 63%. Being overcharged by a mechanic or repair person went from 36% to 52%.
The cause is plain effort. A checklist shows every option at once and nothing makes people consider each one. They tick what jumps out and move on. An empty box can mean "no", or "I didn't read that far". You can't tell which.
Position made it worse. Pew measured an average primacy effect of 3 percentage points for select-all questions, so items near the top got ticked more just for being there. For forced choice it was 0. The post on question order effects explains that mechanism in more detail.
The forced-choice alternative: yes or no per item
Ask each item as its own yes/no question, often laid out as a compact grid. After these results Pew adopted forced choice over select-all in its online surveys whenever possible.
In the past 12 months, has each of these happened to you?
Yes No
Overcharged by a mechanic ○ ○
Lost a job ○ ○
Denied health insurance ○ ○
People have to stop at every row and commit, and a skipped row no longer looks like a "no".
It costs time, and on a phone a wide grid turns into a long scroll. Keep items few, labels short, and test at phone width. The post on survey length and drop-off shows what those extra seconds do to completion.
When a checklist is still good enough
Select-all is a trade-off you can sometimes accept. Pew's data has a useful detail here: the ranking from most to least endorsed was identical or very similar in both formats. The checklist got the levels wrong and mostly got the order right.
That tells you when to use it. If you only need to know which of six integrations people use most, a checklist gives a usable ranking at lower cost. If you need the level itself, say the share of customers who hit a billing error, use forced choice. That is exactly the number a checklist pushes down.
Short, concrete lists suffer least. Twelve vague experiences, each needing a moment of memory search, is where a checklist loses most.
How many options is too many
The bottom of a long list gets less attention than the top. Pew's guidance is to keep answer choices to a relatively small number in most circumstances, "just four or perhaps five at most".
That is a target for opinion items. A list of countries or job titles is long by nature, and a searchable dropdown works better there than forty radio buttons. For lists in between, try this first:
- Split it. Twelve options often hide two questions with six each.
- Group options under short subheadings so people can jump to their block.
- Randomise the order across respondents so position bias spreads evenly.
- Cut options almost nobody picked in your pilot and let "Other" catch them.
Never shuffle a list with a natural order, like frequencies, amounts or age bands. And keep "Other" and "None of these" at the bottom even when the rest moves.
One answer or several: decide by the analysis
"What was the main reason you cancelled?" and "Which of these reasons played a part?" answer different business questions.
A single answer forces a priority. The shares add up to 100% and you get a clear top reason, but secondary reasons vanish. Multiple answers show the full spread with no priority, and the percentages add up to more than 100%, which confuses anyone reading the chart.
Often you want both. Ask which reasons applied, yes or no per reason. Then ask which one mattered most, showing only the reasons they said yes to. The post on conditional logic and feedback shows how to route that.
Don't know, not applicable and prefer not to say are different answers
A missing way out is also an exhaustiveness failure. Someone who never contacted support can't rate it. Without "Not applicable" they skip, if you let them, or they guess, and the guess lands in your satisfaction average.
These options mean different things:
- "Don't know": the question applies, but they lack the information.
- "Not applicable": the question doesn't apply to them.
- "Prefer not to say": they know and won't tell you.
Merge them and you lose the signal: lots of "don't know" on pricing points at communication, lots of "prefer not to say" on income points at trust. The post on asking sensitive questions explains why an explicit way out keeps the other answers honest.
Every visible exit also tempts people who just want to move on, so offer "Don't know" only where not knowing is real and common.
Build the list from real answers, not from a meeting
The surest way to an exhaustive, non-overlapping list is to stop guessing. Ask the question open-ended first, to a small group or in a soft launch, and build the options from what people say. You'll find answers nobody in the planning meeting thought of, and options you separated that respondents see as one.
Coding those answers into categories is covered in analysing open-ended answers. Then pretest the finished list. A cognitive interview is the fastest way to hear "I wasn't sure if I was this one or that one", and the post on pretesting a survey describes how to run one.
A check for every answer list before launch
Go through each closed question with these checks:
- Could any real answer fit two options? Fix the boundaries.
- Is there anyone for whom no option is true? Add it, or put "Other" at the bottom.
- Is it a select-all list where you need the level? Switch to yes/no per item.
- More than five options on an opinion question? Split, group or cut.
- Does the order mean something? If not, randomise and keep the exits at the bottom.
- Do "Don't know", "Not applicable" and "Prefer not to say" appear separately, and only where they're real?
After fielding, monitoring response quality catches careless respondents. It can't recover an answer your list never allowed.
Frequently asked questions
Should I always replace select-all-that-apply with yes/no questions?
Not always. Pew found forced-choice items produce estimates 8 percentage points higher on average, while the ranking of items stayed very similar. If you need the actual share of people who experienced something, use forced choice. If you only need the relative order and the list is short and concrete, a checklist is an acceptable trade-off.
Is an "Other, please specify" option enough to make my list complete?
No. Most respondents choose from the options they see and few type something new, so an Other field undercounts the answers you left out. If a typed answer keeps coming back in the first days, add it as a real option in the next wave.
How many answer options should a multiple-choice question have?
For opinion questions, Pew advises a relatively small number, four or perhaps five at most in most circumstances. Factual questions with naturally long lists, such as countries or job titles, work better as a searchable dropdown or typed answer.
Should I randomise the order of answer options?
Randomise when the options have no natural order, so position effects spread evenly instead of always favouring the first ones on the list. Keep a fixed order for anything with a built-in sequence, such as frequencies, amounts, age bands or rating scales, and keep Other, None of these and Don't know at the bottom.