Back to blog
People mid-stride crossing a white-striped pedestrian crossing, their legs blurred by motion and seen at street level
Survey designAugust 14, 202610 min read

Sample quality: who answers your survey and who doesn't

A big response count is not a representative sample. Here is how coverage, non-response and self-selection quietly bias your results, and what to do about the people your survey never hears from.

By SurveyLane · The team building SurveyLane

You can write flawless questions, pick the right scales, and still end up with data that describes the wrong people. The problem is not the form. It is who filled it in. Every survey pulls answers from a subset of the people you care about, and the gap between that subset and the whole is where quiet bias lives. A clean export of 4,000 responses can mislead you more than 400, because the size hides the fact that the same kind of person answered every time.

A high response count is not a representative sample

The number that makes people relax is the total. "We got 4,000 responses." It feels like enough. But representativeness has nothing to do with how many people answered. It has to do with whether the people who answered resemble the people who didn't. Size buys precision, not accuracy. Skew the sample and more responses just measure the skew more precisely.

This is the most common mistake in in-house survey work. You send a satisfaction survey to your whole list, a few thousand reply, and you report the average as the voice of the customer. It is really the voice of the customer who opens your email and feels strongly enough to click. A real group. Not your customer base.

Response rate versus non-response bias

Response rate is the share of invited people who finished the survey. Worth tracking, but it is not what decides whether your data is trustworthy. That comes down to non-response bias: whether the people who stayed silent differ, on the thing you are measuring, from the people who answered.

The two come apart. Pew Research Center has tracked telephone response rates falling from 36% in 1997 to 9% by 2012, then to 6% by 2018, with contact rates dropping from 65% to 27% over the same stretch. And yet carefully run, properly weighted low-response surveys still produce reasonably accurate estimates. A low response rate is a warning light, not a verdict. The risk is not that few people answered. It is that the few who answered are systematically different from the many who didn't.

Where your sample comes from: the sampling frame

Before anyone can respond, they have to be reachable. The sampling frame is the list of people who could possibly receive your survey: your email list, your active users, the panel you bought, the customers with a phone number on file. Every frame excludes someone, and the exclusions are rarely random.

If your frame is "people with an email in our CRM," you have already dropped every customer who bought through a reseller, paid by invoice, or unsubscribed. If it is "users who opened the app this month," you dropped everyone who churned, which is exactly who a retention survey needs to hear from. Name your frame out loud before you send, and write down who it leaves out. That sentence is the honest boundary of every claim you will make later.

Coverage error: who never had a chance to answer

Coverage error is the mismatch between your frame and the population you actually care about. It is the people your survey could never reach, no matter how good the response rate. The classic version is surveying "the general public" through a channel that leans one way. Web panels under-represent people who are older, less online, or lower-income. An in-app survey misses everyone who deals with you through the website. An in-store intercept misses everyone who shops online.

Coverage error is dangerous because it is invisible from inside the data. You cannot see the people who were never eligible to appear, so compare the frame against what you already know about the population, and say so plainly when a channel cannot reach a group you claim to describe.

Non-response bias: who had the chance but didn't take it

Non-response bias is the more familiar cousin. These people were in your frame and got the invitation. They just didn't answer, and their silence is not random. People who respond tend to be more satisfied or more angry than average, more available, more comfortable in the survey's language, more invested in the topic.

A product survey over-samples power users and the recently burned. An employee survey over-samples people who feel safe enough to speak and people who have a grievance, while the quietly disengaged middle says nothing. The post on spotting low-quality responses deals with bad answers that do arrive. Non-response bias is about the good answers that never do, and it is harder to catch because nothing in the export flags it.

Convenience samples and self-selection

A convenience sample is whoever was easy to reach: a link on your homepage, a poll in your newsletter, a request in a community you already run. Self-selection is the same problem with the volume turned up, because when taking part is fully voluntary and visible, the people who opt in are defined by their motivation to opt in. Caring enough to click is correlated with the very opinions you are trying to measure.

Convenience samples are good for generating hypotheses, checking whether a question is understood, and catching usability problems. They are not good for estimating a proportion in a population or comparing two groups as if the difference were causal. The failure is treating their numbers as if they came from a probability sample. "62% of respondents prefer the new design" is a fact about your respondents. It turns into a lie the moment you drop the word "respondents" and imply it holds for all your users. When you must use an open link, say so, because a self-selected sample and a drawn sample deserve very different levels of confidence.

Margin of error is about sampling, not bias

The margin of error quoted next to a poll describes sampling error and nothing else: the variation you would get from drawing different random samples of the same size from the same frame. It assumes a probability sample and says nothing about coverage error, non-response bias, or a bad question.

So a plus-or-minus three points printed under a self-selected web poll is close to meaningless. It answers how much the number would wobble if you redrew the sample, while the real question is whether these are the right people at all. Quote a margin of error only when you drew a probability sample. If you are computing statistics on scale questions, the guide to designing scale questions covers the measurement side of the same problem.

Weighting: a correction, not a cure

Weighting adjusts your results so under-represented groups count for more and over-represented groups count for less, using known population figures as the target. If your customer base is 55% mobile but only 40% of your respondents are, you can weight mobile responses up to close the gap. Done well, it removes bias on the variables you weight by.

The limits matter as much as the mechanism. Weighting only fixes bias on characteristics you can measure and whose true distribution you know in the population. It does nothing for bias on things you didn't measure, and it cannot invent respondents from a group that produced zero answers. Weight a group that is 2% of your sample up to 20% and a handful of people carry the whole estimate, so the noise explodes even as the bias falls. It is a real tool, not a laundering step that turns a broken sample clean.

Boosting response the honest way

Non-response bias grows when the people who skip differ from the people who answer, so lifting the response rate helps most when it pulls in the reluctant rather than more of the eager. Short surveys finish more often, so cut every question that will not change a decision. A clear invitation from a recognised sender beats a vague one, and deliverability counts here: an invitation in the spam folder is a non-response you caused.

Timing and reminders recover people who meant to answer and forgot. Modest, universal incentives lift response without skewing who answers much, while topic-specific incentives pull in exactly the biased crowd you were trying to balance. An accessible form matters too. As the guide to accessible surveys explains, respondents who can't finish on a screen reader or a phone become silent non-responses that never show up as errors.

Reporting who you missed

The most credible thing in a survey writeup is a clear statement of its limits. It costs nothing, and it is the first thing a careful reader looks for. State the frame, the response rate, and the groups you know are under-represented. If you weighted, say what you weighted by and to what target. If it was a convenience or self-selected sample, label it and keep causal language out.

This is not throat-clearing. A stakeholder who knows the survey only reached engaged users will read a 90% satisfaction score correctly, as "engaged users are happy," instead of stretching it to "customers are happy." The same honesty applies when you feed open-text answers into an AI workflow. The guide to analysing surveys with AI is worth reading alongside this, because a model will happily summarise a skewed sample into confident, wrong themes without ever flagging who was missing.

A pre-send checklist for sample quality

Run through this before the invitation goes out. Name the population you want to describe and the frame you are actually sampling, in one sentence each. List the groups the frame cannot reach, and decide whether you can still make the claim you intend to. Confirm you are pushing the survey to a defined sample rather than posting an open link. Check the form works on a phone and a screen reader, so completion isn't quietly filtered by device. Decide which population figures you could weight by. Then write the limitations paragraph now, before you see the results, so it describes the design honestly instead of defending whatever numbers you happened to get.

Frequently asked questions

Is a low response rate always a problem?

Not on its own. A low rate raises the risk of non-response bias but does not guarantee it. What matters is whether the people who answered differ, on the thing you are measuring, from the people who didn't. Large surveys routinely report single-digit response rates and still produce accurate estimates after weighting. Treat a low rate as a prompt to check for bias, and never treat a high rate as proof the sample is sound.

How large does my sample need to be?

That depends on how precise you need the estimate and how much you plan to slice the data, not on a fixed number. But size is the second question. First decide whether your frame reaches the people you want to describe, because a bigger sample from a biased frame only makes a wrong answer more precise. If you break results down by segment, size each segment rather than the total, since a 4,000-response survey can still hold only 30 people in the group you most need to compare.

Can weighting fix a biased sample?

Only partway. Weighting corrects bias on characteristics you measured and whose true distribution you know, such as region, device, or age. It does nothing for bias on things you didn't measure, and it cannot represent a group that produced no responses. Heavy weighting also inflates the noise, because a few people end up standing in for many. Use it to adjust a decent sample, not to rescue a broken one.

What's the difference between coverage error and non-response bias?

Coverage error is about who could never have answered because they were not in your frame at all, such as customers with no email on file. Non-response bias is about who was in the frame and got the invitation but chose not to answer. You fix them differently: coverage error by broadening the frame, non-response bias by lifting participation among the people who tend to skip.

Further reading