Back to blog
A person in a dark jacket reads through printed pages held above a desk
Survey designAugust 25, 202610 min read

How to pretest a survey before you send it

The cheapest survey fix is the one you make before launch. Here is how to pretest questions with expert review, cognitive interviews and a soft launch so you catch problems while they are still free to fix.

By SurveyLane · The team building SurveyLane

Once a survey is live, every flaw in it is frozen. A confusing question keeps confusing people, a broken skip pattern keeps stranding respondents, and you find out only when the data comes back wrong. Pretesting is where you catch those flaws while they are still cheap to fix, before the first real respondent sees the form.

Why the fix has to happen before launch

A survey is a measuring instrument, and you do not calibrate one after taking the readings. Once responses are in, a badly worded question cannot be repaired. You can drop it and waste the effort every respondent spent on it, or keep it and reason around the flaw, which weakens every conclusion that touches it. Either way, the only real fix happened, or did not, before you hit send.

So pretesting is the best hour you will spend on a survey. An hour watching five people struggle with a draft saves you from analysing a thousand answers to a question they all misread the same way. The problems you catch are ordinary. A word that means two things. A scale that does not fit the question. A required field that blocks people who genuinely have no answer.

The four places a question can break

To pretest well, it helps to know what you are looking for. Survey methodologists model answering a question as four steps, a framework from Roger Tourangeau: the respondent has to comprehend the question, retrieve the relevant information from memory, form a judgement, then map it onto one of your response options. A question can fail at any of the four, and each failure looks different.

Comprehension fails when the wording is ambiguous, uses a term the respondent does not share, or quietly asks two things at once. Retrieval fails when you ask for something people cannot recall: "how many times did you visit our site last month" assumes a memory nobody keeps. Judgement fails when the question demands a summary the respondent has never formed, so they make one up on the spot. Response mapping fails when your options do not contain the answer they want to give, so they pick the closest wrong box. When you pretest, you are watching for the moment a person hits one of these four walls.

Start with expert review, because it is free

The cheapest pretest is another qualified person reading your draft with intent. Not a glance: a real read, question by question, asking of each one what it is measuring and whether a reasonable respondent could read it differently than you mean. Someone who was not in the room when you wrote the survey catches a surprising share of problems, because they arrive without your assumptions.

Expert review is strongest at structural faults: double-barrelled questions, leading wording, missing response options, scales that do not match the construct, skip logic that sends people to the wrong place. It is weakest at how ordinary respondents will actually read a question, because an expert reads too well and resolves ambiguities a first-time reader cannot. Before you write questions at all, the piece on writing better survey questions covers the wording faults expert review is best at catching.

Cognitive interviewing: watch a real person answer

Cognitive interviewing is the core pretest method, and it is simpler than it sounds. You sit with one person from your audience, give them the survey, and have them tell you what is happening in their head as they answer. This is how organisations like the Pew Research Center test new questions before fielding them, and it surfaces problems no amount of expert review will.

There are two techniques, and you use both. The think-aloud asks the respondent to narrate their thoughts as they read and answer: "tell me what you are thinking as you work through this one." You stay quiet and listen. Verbal probing is the follow-up: after they answer, you ask "what did that word mean to you?", "how did you arrive at that number?", or "was there an answer you wanted to give that wasn't there?" The think-aloud shows you where they hesitate; the probes tell you why, and between them they test all four failure points from the last section.

How many interviews, and with whom

You do not need a statistically representative sample to pretest, because you are not measuring anything yet. You are finding faults, and faults repeat quickly. Five to eight cognitive interviews surface most of the serious problems in a typical survey. If three of your first five interviewees stumble on the same item, you do not need a sixth to confirm it.

Who you interview matters more than how many. Pick people who resemble your actual respondents, and deliberately include the ones on the edges: the least expert, the ones taking the survey in a second language, the ones whose situation does not fit your assumptions. A question about "your manager" breaks for someone who has two managers or none, and you only discover that by interviewing them. This is the mirror image of the post on sampling and representativeness: there you want a representative cross-section, here you seek the edges.

Running the session without leading the witness

The failure mode of a cognitive interview is the interviewer rescuing the respondent. The moment someone frowns at a question, the instinct is to explain what you meant. Resist it. The frown is the data; if you clarify, you have tested a version of the question that will not appear in the real survey.

Keep your reactions flat, because respondents read you and adjust: do not nod at answers you like or wince at ones you do not. When someone gives a short answer, "tell me more about that" is almost always the right neutral prompt. Record the session if the person consents, because the exact words a respondent uses for their confusion often show you how to rewrite the question.

Pretest the machine, not just the questions

Questions are not the only thing that breaks; the survey as a running system has its own failure modes. Click through every path, not just the happy one, and if your survey branches, confirm each branch lands where it should. Conditional logic is easy to get subtly wrong, and the post on conditional logic and feedback covers the design side; pretesting is where you verify the wiring holds.

Test the survey on a phone, because many respondents answer on one, and a matrix question that looks fine on a laptop can be unusable at 375 pixels wide. Test with a keyboard only and, if you can, with a screen reader; the post on accessible survey design covers what to check. And confirm required fields are only required where a missing answer is genuinely a problem, because a forced field with no valid option for some people is a wall that ends their session.

The soft launch: a pilot with real respondents

After the qualitative pretesting, run a soft launch: send the survey to a small slice of your real audience, a few percent, and stop to look at the data before releasing the rest. This catches what individual interviews cannot: patterns that only appear at volume. One interview tells you whether a person can answer the question; the pilot tells you what a crowd actually does with it.

Size the pilot to your total. Sending to a hundred thousand people, a few hundred responses is plenty to spot the obvious breakages. If your whole population is three hundred, you cannot spare a useful pilot, so lean on cognitive interviews and treat the first few dozen responses as your check. Either way, you want a moment where the survey has met real respondents but is not yet committed to all of them.

Reading the pilot data for trouble

A pilot is only worth running if you interrogate the results, looking for problems, not findings. Watch the completion rate: if there is a question where people abandon the survey, that question is doing damage. If almost everyone picks the same option, it may not discriminate, or the scale may be pushed against a floor or a ceiling. A question everybody answers identically is measuring nothing.

Check item nonresponse, the share who skip each optional question, because a spike on one item usually means it is confusing, intrusive, or missing the option people wanted. Read the open-text answers, because respondents often use them to tell you the survey itself is broken; the post on analysing open-ended answers covers how to work through them at volume. And watch timing: a completion time far shorter than you expected means people are speeding through without reading, a quality problem covered in the post on monitoring response quality.

When the pilot tells you to change something

If the pilot reveals a real problem, fix it before the main send even though it is tempting not to. You already have responses in hand, and changing the instrument means those responses no longer match. That is why the pilot stays small: a few hundred answers to a flawed question is an acceptable loss, a hundred thousand is not. Treat those responses as pretest data, not part of the clean run. Resist the opposite temptation too, of tuning forever. Past a point you are polishing wording no respondent will notice while the survey sits unsent, and you want no serious faults left, not a perfect one.

A short sequence you can actually run

None of the steps is expensive on its own. Draft the survey, set it aside for a day, self-review against the four failure points, then hand it to a colleague for expert review. Run five to eight cognitive interviews, edge cases included, and rewrite whatever trips more than one person. Click every branch on a phone, with a keyboard, ideally with a screen reader. Soft-launch to a small slice, read the numbers, fix what they expose. Then send the rest. What that saves you from is the one cost you cannot recover: a survey that measured the wrong thing, found out too late to do anything but live with it.

Frequently asked questions

How many people do I need for cognitive interviewing?

Five to eight is enough for most surveys. You are looking for faults, not prevalence, and serious problems repeat quickly, so the same confusing question tends to trip most people who read it. When three of your first five interviewees stumble on the same item, that is already a signal to rewrite it. Interview the right mix of people, especially edge cases, rather than running many sessions.

What is the difference between cognitive interviewing and a soft launch?

They catch different things. Cognitive interviewing is qualitative: you sit with one person and learn why a question confuses them, which aggregate data cannot tell you. A soft launch is quantitative: you send to a small slice of real respondents and read the patterns that only appear at volume, like where people abandon the survey or which question everyone skips. Do the interviews first, then the pilot.

Can I skip pretesting if I am reusing validated questions?

Partly. Validated questions have been pretested for wording, so you can trust the individual items more. But you still have to test the survey as a whole: question order, the skip logic connecting them, how they render on a phone, and whether the combination fits your audience. Validation travels with a question, not with the survey you dropped it into, so a short pilot still helps.

Is expert review enough on its own?

No. Expert review is excellent at structural faults like double-barrelled questions, leading wording and broken logic, and it costs almost nothing, so always do it. But an expert reads too well to see how an ordinary first-time respondent stumbles, because they resolve ambiguities automatically. Covering that blind spot is what cognitive interviewing is for, so you use both.

Further reading