fiction

Howzat for an average? Lessons from cricket and data analysis

Andrew Wiseman looks at why looking at mean scores alone might not unearth all the answers

I had a week off last week. A beautiful, sun-filled week in Northumberland, as it happens. It gave me time to recharge, and to catch up on the test series between England and India. I exchanged messages with a friend about the crazy scorecard from England’s first innings in the 2nd test. For those not that bothered about cricket, England scored 407 runs in their first innings, of which 19 were ‘extras’ (leg byes, byes, no-balls).

Based on the 388 runs scored by batters, this equates to an average of 35.3. But does this tell the whole story? You might have heard the adage of putting your head in the oven and feet in the freezer, where on average you’ll be just the right temperature? Well this innings was exactly like that.

In that 388 total, wicket-keeper Jamie Smith scored 184, whilst Harry Brook also contributed 158. That’s 342 of the total, or 88%. When excluding these two scores, the average for the other nine batters was 5.1.

But back to that old adage of only using mean averages. If we look at both median and mode averages, then the average isn’t 35.3. It’s 0. ZERO. In the innings, six batters scored NO runs, meaning it was the most commonly observed score (mode), and the score in the middle of the series, i.e. the 6th score in the ranked series was zero (median).


The implications for market research

What implications does this have for us in research? Well, we already use the full range of averages when it comes to looking at online survey length. What’s the mean completion time, the median, and the mode. It allows us to look for those outliers that mean that the survey has been completed too quickly, whether fraudulently or without due care by the participant. It’s also useful to look at other averages than the mean if considering some missing data replacement for modelling purposes, where you need complete records. Looking at different averages can help to ensure that the interpolation is as accurate a reflection of the collected data as possible. We might also want to use other averages with any open numeric data, to root out the Jamie Smiths and Harry Brooks, to see if that data is genuine or not.

In short, relying solely on the mean can paint a deceptive picture, one that smooths over the extremes and outliers that often tell the real story. Whether you’re analysing cricket scores or survey data, the lesson is the same: context matters, and so does the distribution beneath the surface.

So next time you see an average, think about what’s really going on underneath? Because sometimes, the truth isn’t in the average, but in the exception.