Rating scales feel simple. Ask a question, collect a score, compare the results. But as Colin Auton explains, the numbers are not always as straightforward as they appear. From choosing between 5, 7 and 10-point scales to comparing results across international markets, small decisions in research design can have a big impact on how research findings are interpreted.
We’ve probably all done it. Market “A” gives us an average score of 7.8, Market “B” gives us 7.4, and we conclude that overall satisfaction is higher in Market “A”.
But is it really?
Rating scales are one of the most familiar measurement tools in market research. We ask people how satisfied they are, how likely they are to buy, how much they agree, and how likely they are to recommend.
They’re simple to answer, easy to analyse and give clients numbers they can quickly get their heads around. But there’s an assumption sitting underneath all of this: that everyone uses a rating scale in broadly the same way.
Unfortunately, people don’t.
Even within one market, some people are naturally more willing to give extreme scores. Others rarely venture beyond the middle of the scale. And while one person might agonise over whether to give a six or a seven rating, another might see very little difference between the two or refuse to ever give the highest score.
Take that across different countries, languages and cultures and things become even more interesting.
So, when we compare that 7.8 with a 7.4, are we measuring a genuine difference in opinion? Or, at least in part, a difference in how people use the scale?
Five, seven or ten. Does more really mean better?
Let’s start with the recurring question: how many points should a rating scale have?
There isn’t a perfect answer. Sorry.
For most general consumer research, a five-point scale is a pretty good place to start. It gives people enough room to express differences in opinion without asking them to make distinctions they may not actually be capable of making.
A seven-point scale can be useful where we genuinely need more nuance, particularly when people are engaged with the subject and likely to have a well-formed view. Also, for when we’re needing to work the data hard to find the meaningful differentiation in a segmentation.
Then there are 1 to 10 and 0 to 10 scales. They’re familiar, intuitive and have an obvious role where there’s an established convention, such as likelihood to recommend.

But more points don’t automatically mean more insight
It’s tempting to think of a scale like a ruler. If five points gives us one level of accuracy, surely ten gives us twice as much? Human judgement doesn’t work quite like that. Respondents aren’t usually making a precise calculation in their heads. They’re looking at the options and deciding which one feels about right.
So, can somebody meaningfully distinguish between giving something a six rather than a seven? Probably. Can they consistently distinguish between every adjacent point on a ten-point or eleven-point scale across a 15-minute questionnaire? I’m less convinced.
This matters more when we analyse the results.
If an average satisfaction score moves from 7.2 to 7.4, the decimal point makes it look reassuringly precise. But we shouldn’t confuse more precise numbers with more precise opinions.
My view?
Use as many scale points as you need. Not as many as you can.
Should we let people sit on the fence?
Then there’s the midpoint. Some people don’t like them. The argument goes that if you remove the middle option, respondents have to make their minds up. Are they positive or negative? Satisfied or dissatisfied? For or against? It can be particularly tempting when a client wants a decisive answer.
The problem is that if somebody is genuinely neutral, forcing them to choose doesn’t reveal their opinion. It creates one. For most attitudinal questions, I favour a midpoint where being neutral genuinely makes sense. If someone is neither satisfied nor dissatisfied, let them tell us that.
The more important issue is not to make the midpoint work too hard. ‘Neither agree nor disagree’ is not the same as ‘Don’t know’. And ‘Don’t know’ is not the same as ‘I’ve got no experience of this’. If all three are legitimate answers, we should give respondents a way to tell us the difference.
There are occasions when a forced choice makes sense. But it should be because the research question genuinely requires someone to choose a side, not because we’ve decided sitting on the fence is inconvenient.

The bigger problem: people don’t use scales in the same way everywhere
This is where international research gets particularly interesting.
Give people exactly the same scale in different markets and you won’t necessarily get directly comparable behaviour.
Research into cross cultural response styles has identified differences in the tendency to use extreme scores, middle categories and agreement-based responses. Those differences can affect international comparisons even when the questionnaire itself has been kept consistent.
Some respondents are more comfortable using the very top and bottom of a scale. Others tend to stay closer to the middle. In some markets, there can also be a greater tendency to agree with statements regardless of their exact content.
Does that mean we can say “people from Country “X” always do this” and “people from Country “Y” always do that”? Definitely not.
There’s a danger of taking broad cultural patterns and turning them into fairly lazy stereotypes. How somebody responds to a questionnaire can be influenced by their age, education, language, the topic, how engaged they are, how the survey is being administered and plenty more besides.
But the bigger point still stands.
A four-point difference between two markets doesn’t automatically represent a four point difference in what people actually think.
There are some really practical differences to think about too.
We’re used to the idea that a higher score is a better score. But that isn’t a universal convention. German educational grading is one obvious example, where 1 represents the strongest performance.
That doesn’t mean a German respondent is suddenly going to interpret a clearly labelled 0 to 10 satisfaction scale backwards. It DOES mean we shouldn’t assume that the numbers explain themselves. If 0 means ‘Not at all satisfied’ and 10 means ‘Completely satisfied’, then say so.

Lost in translation?
We tend to spend a lot of time making sure the question itself has been translated properly in international research. Quite right too.
But what about the scale?
Words such as ‘fairly’, ‘quite’, ‘somewhat’, ‘very’ and ‘extremely’ can be surprisingly tricky. The literal translation may be correct, but does the strength of the word feel the same?
If ‘fairly satisfied’ and ‘very satisfied’ feel like two meaningfully different positions in the UK, we want the translated versions to create roughly the same distinction elsewhere.
So the aim isn’t just to translate the words. It’s to translate the meaning of the scale.
Presentation matters too. The way a scale is displayed needs to make sense in the language and market in which it’s being used, while the underlying coding and meaning stays consistent.
These can feel like fairly minor questionnaire design details. They feel much less minor when somebody puts the results from 12 countries into a league table and starts asking why France is three points behind Spain.
So what should we do with international results?
I’m definitely not suggesting we stop comparing markets. That would make quite a few of our projects difficult!
But I do think we need to be more thoughtful about what those comparisons are actually telling us.
1. Don’t just look at the average
A mean score of 7.0 can hide a multitude of sins.
It could mean nearly everyone gave you a seven.
Or half your respondents gave you a ten and the other half gave you a four.
Same average. Completely different story.
Look at the distribution. Look at how many people are using the extremes. Look at the midpoint. Look at top box or top two box scores where they’re useful.
There is nearly always more going on than the average tells you
2. Don’t get too excited about tiny differences between countries
If one market scores 7.6 and another scores 7.3, we should be cautious about immediately declaring the first one more positive.
It might be.
But part of that difference could reflect how the two groups use the scale.
Statistical significance is important, but it can’t tell us whether everybody interpreted and used that scale in precisely the same way.
3. Look for patterns within and across markets
Often, the relative story is more useful.
If Brand “A” beats Brand “B” in the UK, France and Germany, that tells us something quite interesting even if the absolute scores differ between countries.
Likewise with tracking. If the same measure moves consistently within a market over time, that can be more useful than obsessing over whether one country’s absolute score is a few points higher than another’s.
4. Keep the things you CAN control consistent
There are already enough potential sources of difference, so don’t introduce more unnecessarily.
Keep the number of scale points consistent. Keep the meaning of the anchors consistent. Keep question wording and survey mode as comparable as possible. Treat ‘Don’t know’ in the same way.
Localise where you need to. But changing a five point scale to a ten point scale in one country simply because somebody thinks the locals prefer scores out of ten probably creates more problems than it solves.
So, which scale should we use?
After all that, you probably want an answer.
For most general consumer research, I’d be pretty comfortable starting with a fully labelled five-point scale.
Seven points can be useful when there’s a genuine need for more nuance.
0 to 10 makes sense when there’s an established reason for using it, or when people naturally think about the question that way.
And I’d generally include a midpoint when being genuinely neutral is a valid response, while keeping ‘Don’t know’ or ‘Not applicable’ separate where they’re needed.
But the bigger lesson isn’t really about whether five beats seven.
It’s about remembering that a rating scale isn’t a perfectly neutral ruler.
You can standardise the questionnaire. You can use exactly the same numbers. You can translate every question carefully. And you can still end up with people using that scale in subtly different ways.
So by all means compare markets.
Just look beyond the average, be wary of false precision and don’t assume that because two respondents both clicked seven, they necessarily meant exactly the same thing.
Sometimes, a seven isn’t just a seven.
Rating scales are widely used in market research to measure satisfaction, agreement, purchase intent and recommendation. However, scores are influenced by how people interpret and use scales, meaning comparisons across audiences and countries require careful analysis. Effective questionnaire design considers scale length, wording, cultural differences and the decision the research needs to support.
A survey score is only meaningful if you understand how it has been created. Choosing the right scale, designing questions carefully and interpreting results in context are all essential to making confident decisions from research.
At Mustard, we help organisations design research that delivers clarity rather than just numbers. From questionnaire design and international research to interpreting complex datasets, we make sure the insight behind the score is what drives decisions.
If you’re planning customer research and want to make sure your survey design delivers meaningful answers, get in touch.
