Survey Response Scales: How to Choose the Right One for your Questionnaire

Survey scales

How you design a survey can change the answers you get. Question wording, question order, response labels, the number of response options, and whether you include a midpoint can all influence how respondents answer.

A response scale might use numbers such as 1–5 or 1–7, verbal labels, opposing adjectives, or a binary choice such as Yes/No.

There is no universal “best” survey scale. The right format depends on what you are trying to measure, who is answering, how the survey is administered, and how you intend to analyze the results.

TL;DR

  • Match the response scale to the thing you are actually trying to measure.
  • Five- and seven-point scales are common for attitudes, but there is no universally optimal number of response options.
  • Binary questions work when the underlying decision is genuinely binary, but they remove nuance.
  • Individual rating items are usually ordinal; don’t assume the distance between every response category is equal.
  • Small changes to labels, response order, midpoints, and question wording can materially change survey results.
  • Pretest important surveys before using the data to make major decisions.

Start with the question, not the scale

Before choosing a response scale, define exactly what you need to learn.

Are you measuring satisfaction, frequency, agreement, perceived difficulty, likelihood, preference, or an observable behavior? Different constructs call for different response formats.

Question wording also matters as much as the scale itself. Response options should be mutually exclusive, cover the reasonable answers a respondent might give, and use language your audience understands.

For important surveys, pretest new questions with people similar to your target respondents. Pew Research Center and AAPOR both recommend cognitive testing, pilot testing, or other forms of pretesting before fielding a questionnaire.

How to design survey rating scales

Surveys are a great source of insight into your visitors’ attitudes. Certain surveys also allow you to compare yourself with the competition. But as Jared Spool notes, the nuances of survey scale design add challenges:

“My point with this example is that scale design and anchor choice will influence respondents’ ratings—both higher and lower.

This is a key reason why I’m skeptical of the cross-company comparison data sets where each company is using a different survey instrument. So many variables are in play that legitimate comparisons are quixotic.

Jared Spool

So what can you do to get accurate data? It starts with understanding some of the differences and shortcomings of survey scales.

3 types of survey response scales

When designing surveys, there are many ways to structure survey responses. Three common formats are:

  • Dichotomous;
  • Rating scales;
  • Semantic differential scales.

We’ll look at each, with some survey scale examples to help you visualize them.

1. Dichotomous scales

Dichotomous scales have two choices that are diametrically opposed to each other. Some examples:

  • “Yes” or “No”.
  • “True” or “False”.
  • “Fair” or “Unfair”.
  • “Agree” or “Disagree”.

(Image source)

Dichotomous questions deliberately reduce the answer to two alternatives. That can be useful when the underlying decision is genuinely binary, but it removes the ability to express degrees of opinion or uncertainty.

A midpoint or central-tendency bias is not relevant to a two-option scale because there is no middle category. With longer questionnaires, the larger concern is satisficing: respondents may stop evaluating each question carefully and choose answers that require less effort.

Be particularly careful with Agree/Disagree formats. Survey research has repeatedly found acquiescence bias, where some respondents are more likely to agree with a statement regardless of its content. Pew therefore often prefers questions that ask respondents to choose between alternative positions rather than simply agree or disagree with a statement.

2. Semantic differential scales for questionnaires

So even if satisfied and dissatisfied are “common practices,” they may not be “best practices”—especially in user experience research. You’re trying to delight customers, not just “satisfy” them.

Semantic differential scales gather data and “interpret based on the connotative meaning of the respondent’s answer.” These scales usually have dichotomous words at either end of the spectrum.

They measure more specific attitudinal responses. An earlier AllTrails survey is an example of this:

The number of scale points still requires judgment. Research does not support one universally optimal length, although five- and seven-point formats are commonly used for attitudinal measures.

3. Rating scales

You’re probably most familiar with rating scales (e.g., “On a scale of 1–10, how satisfied were you with our service today?”). The three most common rating scales are:

  • 1–10
  • 1–7
  • 1–5.

The number of response options can affect reliability, discrimination, and how easily respondents can express their opinion.

But there is no rule that a larger scale is automatically worse, or that every survey should use five points.

A 2026 review of survey-scale research found that five- and seven-point formats remain the most common, that scales with very few categories can lose information, and that improvements tend to level off once scales reach roughly six or seven options.

A study comparing three-, five-, and seven-point Likert-type scales similarly found that five points produced better reliability than three, while providing little disadvantage compared with seven in that particular dataset.

The choice depends on your audience and the precision they can provide. More response categories only help if respondents can reliably distinguish between them.

Likert scales for surveys

Likert-type items commonly use five or seven ordered response options, often ranging from Strongly Disagree to Strongly Agree. Other scale lengths exist, but a generic 1–10 rating question is not automatically a Likert scale simply because it uses numbers.

Another great point from Spool’s talk touches on Likert Scales. He rails against the labels we use on scales (satisfied and dissatisfied) instead of the scale itself:

It’s about how we create the scale. We start with this neutral point in our scales. This is how a five-point Likert scale works.

We add two forms to it—in this case, satisfied and dissatisfied—and then, because we think that people can’t just be satisfied or dissatisfied, we’re going to enhance those with adjectives that say “somewhat” or “extremely.”

Okay, but extremely satisfied is like extremely edible. It’s not that meaningful a term.

What if we made that the neutral, and we built a scale around delight and frustration? Now we’ve got something to work with here. Now we’ve got something that tells us a lot more.

We should not be doing satisfaction surveys; we should be doing delight surveys. We need to change our language at its core to make sure we’re focusing on the right thing. Otherwise, we get crap like this.

Jared Spool

More scale points should not be confused with more meaningful precision. A respondent may be able to distinguish clearly between five or seven levels of satisfaction while finding the difference between, say, a 7 and an 8 on a ten-point scale much less interpretable.

Current methodological reviews generally find that five- and seven-point scales offer a practical balance between discrimination and usability, although the best choice still depends on the construct, audience, and analysis.

Which survey scale should you use?

After all these examples of survey rating scales, you might be wondering which to use. They all look useful. However, it depends on the type of data you want, and what survey questions you’re asking. Good survey questions will help inform your rating scale choice.

Dichotomous scales (“yes” or “no”) are great for precise data, but they don’t allow for nuance in respondents’ answers. For instance, asking if a customer was happy with an experience (yes or no), gives you almost no insight into how to improve the experience.

A multi-point rating scale can capture degrees of opinion that a Yes/No question cannot. But the scale should still match the construct being measured. Although—and this is a big point—says Spool, “Anytime you’re enlarging the scale to see higher-resolution data, it’s probably a flag that the data means nothing.”

NPS is a specific standardized case. Its recommendation question uses a 0–10 scale, with respondents grouped into promoters, passives, and detractors. If you want a conventional NPS score that is comparable over time, changing that response scale changes the metric itself.

For behavioral questions, it is often better to ask about concrete frequency or quantity than to convert the behavior into a vague agreement scale. For attitudes and perceptions, five- or seven-point ordered scales are common choices.

Likert scales (satisfied vs. dissatisfied) are a little generic for attitudes.

There’s also an older scale, the Guttman scale, that puts a twist on dichotomous and Likert scales. You ask a series of questions that build on each other and escalate in intensity. Here’s a great example from changingminds.org:

Spool talked about the Guttman scale in its relation to customer surveys, saying:

If you’re not happy enough to recommend the product, you’re not going to be confident, and you’re not going to feel it has good integrity if you’re not confident, and you’re not going to have pride in it unless they have good integrity, and you’re definitely not going to be passionate about them unless they do everything else.

This can be a useful tool for measuring satisfaction.

Ordinal and interval scales

Developed by S.S. Stevens and published in a 1946 paper, there are four types of these scales:

  1. Nominal;
  2. Ordinal;
  3. Interval;
  4. Ratio.

There’s perpetual debate about ordinal and interval scales.

Ordinal scales are numbers that have an order, like “a runner’s finishing place in a race, the rank of a sports team, and the values you get from rating scales used in surveys or questionnaires like the Single Ease Question.”

With ordinal scales, if you’re asking a customer how satisfied they were on a scale of 1–5, a 4 doesn’t necessarily mean they were twice as satisfied as a 2. The difference between a 1 and a 2 isn’t necessarily the same as the difference between a 4 and a 5.

Interval scales establish equal distances between ordinal numbers—for example, when we measure temperature in Fahrenheit. The difference between 19 and 20 degrees is the same as between 80 and 81.

According to Jeff Sauro, Founder of MeasuringU, rating scales can be scaled to be interval:

Rating scales can be scaled to have equal intervals. For example, the Subjective Mental Effort Questionnaire (SMEQ) has values that correspond to the appropriate labels.

You can see the distance between the numbers is equal, but the labels vary depending on how people interpreted their meaning. (translated from Dutch)

Jeff Sauro

What’s the practical difference?

An individual survey rating such as 1 = Very dissatisfied through 5 = Very satisfied is ordinarily treated as ordinal data. You know the order of the categories, but you cannot automatically assume that the psychological distance between 1 and 2 is identical to the distance between 4 and 5.

That matters because calculating a mean implicitly treats the numerical gaps as meaningful.

In practice, the analysis of rating-scale data is more nuanced. Researchers frequently combine several related Likert-type items into a composite score and analyze that score using parametric methods. Research has also shown that many common parametric tests are reasonably robust to departures from strict interval assumptions.

That does not make the ordinal/interval distinction irrelevant. A review of ordinal response scales emphasizes that converting ordered labels into numbers does not prove that respondents perceive equal distances between those labels.

Does it matter whether your data is interval or ordinal?

Yes, because the measurement assumption affects what your statistics mean.

For an individual rating item, showing the distribution of responses, percentages, median, or mode often communicates more than reporting an average to several decimal places.

For multi-item scales or larger analyses, means and parametric methods may still be defensible depending on how the instrument was constructed, the distribution of the data, the sample size, and the analysis you are running.

The important point is to make the assumption explicit rather than automatically treating every numbered survey response as continuous data.

The limitations of survey scales

Even if you design the perfect survey with the appropriate scales, there are limitations. This is especially true if you run a limited range of surveys or conduct surveys sporadically (and without other forms of conversion research).

The meaning behind the numbers

When you run a metric such as NPS, you get a number that can be tracked over time. Cross-company comparisons require more caution because differences in sample, timing, audience, and survey administration can affect the result.

I haven’t heard a better explanation than from Spool’s talk on design and metrics:

I was so disappointed when the people at Medium sent me this: ‘How likely are you to recommend writing on Medium to a friend or colleague?’ It’s not even a 10-point scale. It’s an 11-point scale, because 10 was not big enough.

This is called a Net Promoter Score, and with Net Promoter Scores, if you look at the industry averages that everybody wants to compare themselves to, the low end is typically in the mid-60s and the high end is typically in the mid-80s.

You need a 10-point scale because, if you had a three-point scale, you could never see a difference. Anytime you’re enlarging the scale to see higher-resolution data, it’s probably a flag that the data means nothing.

Here’s the deal. Would a net promoter score for a company like, say, United, catch this problem?

Alton Brown bought a $50 guest pass to the United Club in L.A. and had to sit on the floor. I wonder what his Net Promoter Score for that purchase would be? It probably wouldn’t tell anybody at United what the problem is.

But that’s a negative. What about the positive side?

What’s actually working well? Customers of Harley-Davidson are fond of Harley-Davidson, so fond that they actually tattoo the company’s logo on their body. This is branding in the most primal of definitions.

Jared Spool

Even the Net Promoter System pairs the 0–10 recommendation question with a follow-up asking why the respondent gave that score. The number helps summarize the response; the follow-up provides diagnostic context.

It also depends what you’re selling. As Caroline Jarrett, author of Forms That Work, said:

Just from the point of view of using a Net Promoter Score as a question in a survey, we have to ask whether that question means as much to the people answering it as it might to the business.

There are some things where ‘I’ll recommend this to a friend’ is a really important thing that people would actually do. But there are other things where you’d never recommend it to a friend because you don’t do recommending, and you certainly don’t do recommending of those type of things.

So you might actually be very enthusiastic about the product, but you just might not ever feel the urge to recommend hemorrhoid cream to your pals, you know?

That’s not then giving a true measure of the value of that product. I have my skepticism about Net Promoter Score.

Caroline Jarrett

All of this is to say that ratings scales can tell you a lot, but they can’t tell you everything. Be skeptical when people tell you there’s one question that will tell you how your company is doing.

Little tweaks, big differences

Almost any factor can influence the outcome of a survey, which is why Spool highlights the difficulty of accurate benchmarking data.

GreatBrook, a research consulting firm, ran an experiment with a client in which they created a bunch of surveys with the same attributes, just different scale designs. They gave the questionnaires to 10,000 people and found some interesting things:

  • Providing a numeric scale with anchors only for the endpoints (i.e. a 1–5 scale presented verbal descriptions for only the 1 and 5 endpoints) led more people to choose the endpoints.
  • Presenting a scale as a series of verbal descriptions (e.g., “Are you extremely satisfied, very satisfied, somewhat satisfied, somewhat dissatisfied, very dissatisfied, or extremely dissatisfied?”) led to more dispersion and less clustering of responses.
  • A “school grade” scale led to even more dispersion. A school grade scale asks the respondent to grade performance on an A, B, C, D, and F scale.

Current survey guidance reaches the same broader conclusion. Pew Research Center notes that the number, wording, and order of response options can all affect answers. For unordered lists, randomizing options can help distribute order effects; ordered scales should normally stay in their logical sequence.

Using appropriate language and scales

For certain information (e.g., age), there are many ways you can ask for it. Each produces a different level of precision.

According to MyMarketResearchMethods.com, if you need to calculate an exact average age, ask for age as a numerical value rather than grouping respondents into age bands.

Asking for an exact age gives you ratio data. Grouping age into bands such as 25–34 and 35–44 gives you ordinal categories: the categories have an order, but you lose the exact value.

Neither format is universally better. Exact values provide more analytical precision, while ranges can reduce the sensitivity or effort involved in answering some demographic questions.

For sensitive information such as income, respondents should also have a reasonable way to decline. AAPOR recommends allowing people to skip difficult or sensitive questions or providing an explicit “prefer not to answer” option rather than forcing a response.

Grouping survey responses based on known characteristics

If you use ranges, choose boundaries that make sense for the population and the decisions you need to make.

Income bands appropriate for students may be useless for senior executives. The useful brackets may also change by country, currency, household type, and whether you are asking about personal or household income.

Avoid creating ranges simply because they look evenly spaced. Use categories that reflect the audience or a recognized external classification, and make sure every response can fit into one—and only one—category.

As Balon advises, “For whatever population you’re studying, make sure those income breaks line up with the known characteristics of the population. Not doing this can create additional bias.”

Similarly, writing surveys in your customers’ own language is important. Use the phrases, jargon, and emotions that your customers are familiar with.

How do you do that? You get on the phone and talk to your customers. Or run focus groups. Or run some on-site surveys.

Best practices for demographic insights

With sensitive information like demographic info, how do you establish which defaults to use, which words to use, which scale to use, etc.?

Other than focus groups and interviews, there are some general guidelines and best practices (listed here). If you follow these, your respondents will likely have taken surveys like it before and, therefore, will know how to answer questions based on past experience.

For example, one common way to group adult age is:

  • Under 18 years;
  • 18 to 24 years;
  • 25 to 34 years;
  • 35 to 44 years;
  • 45 to 54 years;
  • 55 to 64 years;
  • 65 or older.

The right categories depend on why you need the information. Only ask for demographic variables that you genuinely intend to use, and avoid forcing answers to sensitive questions.

Response options should be exhaustive and non-overlapping. For sensitive items, including “prefer not to answer” can reduce the pressure to provide an inaccurate response or abandon the survey.

You can also have an experienced market-research consultant come in and tell you if you’re running things well. But, of course, that’s a another expense.

Ultimately, you have to balance the level of specificity you want with the comfort level of your audience.

Frequently asked questions about survey response scales

What is a survey response scale?

A survey response scale is the set of options respondents use to answer a closed-ended survey question.

Is a 5-point or 7-point scale better?

Neither is universally better. Both are widely used, and the choice depends on the construct, audience, and level of useful discrimination.

How many points should a Likert scale have?

Five and seven are common choices, with research generally finding diminishing benefits as scales become much longer.

Should a survey scale include a neutral midpoint?

Include one when neutrality is a meaningful response; removing it deliberately forces respondents toward one side of the scale.

Are Likert responses ordinal or interval data?

Individual Likert-type responses are generally ordinal, although composite scales are sometimes analyzed as approximately interval data.

Is NPS a Likert scale?

No. NPS uses a standardized 0–10 recommendation rating and a specific scoring method.

Should demographic questions use ranges?

Use ranges when they provide enough precision and make the question easier or less sensitive to answer; use exact values when the analysis genuinely requires them.

Key takeaways on survey response scales

The response scale is part of the measurement, not just the interface.

Changing the number of options, labels, midpoint, direction, or wording can change the distribution of answers you receive. That makes scale design especially important when you want to compare groups or track a metric over time.

Start with the construct you need to measure. Choose a response format that respondents can meaningfully distinguish. Keep categories mutually exclusive and sufficiently complete. And do not create extra numerical precision that the underlying judgment cannot support.

Most importantly, pretest important surveys. A response scale that looks obvious to the person who wrote it may not be interpreted the same way by the people answering it.

Improve your customer research

Survey scales are only one part of collecting useful customer evidence. The quality of the question, sample, research method, and analysis matters just as much as the response format.

Build stronger user research skills: CXL’s User Research course covers qualitative and quantitative research, recruiting, interviews, and interpreting user evidence.

Use surveys to generate better experiments: The Strategic Marketing Experimentation course includes customer surveys, on-site polls, usability research, and turning research into test hypotheses.

Turn customer language into marketing insight: The Voice of Customer Data course covers gathering, analyzing, and applying customer language and feedback.

Related Posts

Current article:

Survey Response Scales: How to Choose the Right One for your Questionnaire

Categories

Become an AI native marketer

A six-week live cohort for marketers. Every week you build one AI native workflow you can use at work: a 90-minute live workshop on Tuesday, then you build it on your own data with us in the community, and demo it on Friday.

You don't need more AI tips. You need five working AI workflows. Starts 28 September.

See the program