We love a simple number. One figure that can wrap up a spreadsheet, a trend, a person, a policy. The average is the tool we grab first. It feels fair, sanding down the extremes into a comfortable middle. But that comfort is often a statistical mirage, and when you force it onto data that resists being averaged, you don’t get clarity. You get distortion. The trouble isn’t the arithmetic. It’s the assumption that every dataset has a meaningful center.
Jerome Leland has spent decades watching how numbers get used and misused in public conversation. The mistake isn’t always malicious. Sometimes it comes from rushing, sometimes from a wish to make complicated things digestible for an audience with a short attention span. But the consequences are real. When we average what should not be averaged, we erase the exact information we need to make sound decisions.
The Arithmetic Mean: A Tool With Limits
The arithmetic mean—sum the values, divide by the count—gets taught early and used everywhere. It’s the perfect summary for data that’s symmetric and unimodal, like the heights of adult men in a specific population, or the manufacturing tolerances of a well-calibrated machine. In those cases, the average sits near the center of a bell-shaped distribution, and most individual values cluster around it. The average is representative.
But the world is full of data that isn’t symmetric. Income distributions, for example, are heavily right-skewed. A small number of extremely high earners yank the arithmetic mean upward, far away from what a typical person actually earns. Reporting the average income without disclosing the median—the point where half earn more and half earn less—is a classic, and often deliberate, statistical misdirection. The average becomes a tool for hiding inequality rather than revealing a central tendency.
This isn’t a subtle point, yet it gets routinely ignored in headlines and political speeches. The problem isn’t that the average is mathematically wrong; it’s that it answers a question nobody asked. When people hear “average,” they imagine the typical case. For skewed data, the arithmetic mean describes no actual case at all.
Ordinal Data: When Numbers Are Just Labels
An even more basic error happens when we average data that isn’t truly numerical. Ordinal data—rankings, ratings, categories assigned numbers for convenience—carries order but not magnitude. The difference between a 1-star and a 2-star rating is not the same as the difference between a 4-star and a 5-star rating. The numbers are placeholders, not measurements.
Consider a customer satisfaction survey where respondents choose from “Very Dissatisfied” (coded as 1), “Dissatisfied” (2), “Neutral” (3), “Satisfied” (4), and “Very Satisfied” (5). Calculating a mean score of 3.7 is common practice, but it’s statistically indefensible. The result implies a continuous scale with equal intervals, which doesn’t exist. A score of 3.7 doesn’t mean the average respondent is slightly more than “Neutral.” It’s a numerical ghost, a fiction created by treating labels as measurements.
The proper summary for ordinal data is the median or the mode, and the distribution should be reported in full. A product with fifty 5-star ratings and fifty 1-star ratings has an average of 3, but that average conceals a deeply polarized reception. The mean tells a story of mediocrity; the distribution tells a story of controversy. Which story is more useful? Almost always, the latter.
Bimodal and Multimodal Distributions: The Average That Describes Nobody
Some datasets have two or more distinct peaks. The classic example is human height when sexes are combined. Adult men and women have overlapping but distinctly different height distributions. The overall average height for a mixed-gender population falls somewhere between the male and female averages, and it describes almost no one. It’s a statistical artifact, not a human being.
This problem appears in more subtle forms across media reporting. Consider a study on screen time that lumps together children under five, school-age children, and teenagers. The average screen time might be reported as four hours per day, but that number masks enormous variation. A toddler watching educational videos for thirty minutes gets averaged with a teenager spending eight hours on social media. The resulting figure is useless for policy, useless for parenting advice, and actively misleading for public understanding.
When data is multimodal, the responsible approach is to report the modes separately. Describe the clusters. Acknowledge that the population isn’t homogeneous. The average, in these cases, is a lie of aggregation.

Time-Series Data and the Ecological Fallacy
Averaging over time introduces its own species of error. A stock price that fluctuates wildly between $10 and $100 may have an average of $55, but that number tells you nothing about the risk an investor faced. The sequence matters. The volatility matters. The average smooths away the very information that drives decisions.
In climate science, this error has profound implications. Global average temperature is a carefully constructed metric, but it’s often misinterpreted. A rise of 1.5°C in the global average doesn’t mean every place gets 1.5°C warmer. Some regions warm much more; others may temporarily cool. The average conceals regional extremes, shifts in precipitation, and changes in seasonality. When a policymaker says “we can adapt to 1.5 degrees,” they’re adapting to a number, not to the lived reality of heatwaves, floods, and crop failures that the average hides.
The ecological fallacy—assuming that group-level statistics apply to individuals within the group—is a constant companion to improper averaging. A county with an average income of $80,000 may contain neighborhoods of deep poverty and enclaves of extreme wealth. The average is true for the county as a unit, but false for nearly every person in it.
Proportions and Percentages: The Unweighted Mean Trap
Another common blunder is taking the average of percentages without weighting by the base. If a small hospital has a 100% patient satisfaction rate from 10 patients, and a large hospital has an 80% rate from 1,000 patients, the unweighted average is 90%. This suggests a level of satisfaction that isn’t remotely representative of the patient experience. The vast majority of patients were 80% satisfied, but the small hospital’s perfect score—based on a tiny sample—drags the average upward.
This error appears routinely in school district comparisons, product reviews, and medical success rates. The fix is simple: always weight by the denominator. But the simple fix is often ignored because the unweighted average is easier to compute and, sometimes, because it produces a more favorable number for the party doing the reporting.
When Averages Become Dangerous: Policy and Medicine
The consequences of averaging inappropriate data move from misleading to dangerous when they inform policy or medical decisions. Drug trials that report average efficacy may conceal the fact that a medication works well for some patients, has no effect on others, and harms a small subset. The average benefit looks positive, but the individual experience is a lottery. Personalized medicine exists precisely because averages fail to capture this heterogeneity.
In public health, averaging infection rates across a large geographic area can hide hotspots that need immediate intervention. During the early months of the COVID-19 pandemic, national averages obscured the catastrophic situation in specific cities and nursing homes. Decision-makers who relied on the average delayed action; those who looked at the distribution acted faster and saved lives.
Economic policy suffers from the same flaw. Average wage growth can be positive while the median wage stagnates or falls, if gains are concentrated at the top. Reporting the average without the median gives citizens and lawmakers a false sense of progress. The average becomes a tool for maintaining the status quo, not for measuring the well-being of the typical person.

When the Average Is Not Even a Number
Some data cannot be averaged because the result is not just misleading—it’s meaningless. Consider categorical data like marital status, blood type, or political party affiliation. Assigning numbers to these categories and computing a mean produces absurdity. The average blood type is not a useful concept. The average political party is not a moderate; it’s a statistical ghost.
Yet this kind of averaging creeps into survey reporting. Researchers assign numerical codes to responses for convenience in data entry, then forget that the codes are not quantities. A mean of 2.4 for religious affiliation (where 1=Christian, 2=Muslim, 3=Jewish, etc.) is not a person who is slightly more Muslim than Christian. It’s a category error, a confusion of measurement scales that renders the result uninterpretable.
Nominal data—categories without order—must never be averaged. Ordinal data—rankings without equal intervals—should not be averaged, though the median may sometimes be acceptable. Only interval and ratio data, where differences are meaningful and a true zero exists, can be properly averaged. Even then, the distribution must be checked for skew and outliers.
The Simpson’s Paradox: When Averages Reverse Truth
Simpson’s paradox is one of the most startling demonstrations of how averaging can invert reality. It occurs when a trend appears in several different groups of data but disappears or reverses when these groups are combined. The classic example involves university admissions: within each department, women may have a higher acceptance rate than men, but overall, men have a higher acceptance rate because women apply in larger numbers to departments with lower acceptance rates. The aggregate average tells the opposite story of the disaggregated data.
This paradox appears in medical studies, social science research, and business analytics. A marketing campaign may show improved conversion rates in every segment, yet the overall conversion rate drops because the mix of segments shifted toward lower-converting groups. Averaging across segments without accounting for composition changes produces a false narrative of decline.
The only defense against Simpson’s paradox is to examine data at the appropriate level of granularity. Averages computed across heterogeneous groups are not summaries; they are distortions.
Practical Guidance: When to Avoid the Average
For anyone working with data—journalists, analysts, policymakers, or curious citizens—a few simple checks can prevent the misuse of averages. Before computing a mean, ask these questions:
- Is the distribution symmetric? Plot a histogram. If the data is skewed, the mean will be pulled toward the tail and will not represent the typical value. Use the median or report the full distribution.
- Are there multiple peaks? If the data clusters into distinct groups, the mean falls in an unpopulated valley. Describe the clusters separately.
- Is the data ordinal or nominal? If the numbers are just labels, do not compute a mean. Use the mode or report category frequencies.
- Are the groups of different sizes? If you are averaging percentages or rates from subgroups, weight by the subgroup size. An unweighted mean of percentages is almost always wrong.
- Does the average answer the question you are actually asking? If you want to know what is typical, the mean may not tell you. If you want to know the total impact, the mean multiplied by the count may be appropriate, but only if the data is ratio-scale and the distribution is understood.
These checks take seconds. They prevent errors that can last for years.

The Median: A Better Default
For most real-world data that reaches the public, the median is a more honest summary than the mean. The median is resistant to outliers, works for ordinal data, and always represents an actual data point when the sample size is odd. It doesn’t create fictional values. When a journalist reports that the median home price in a city is $350,000, that number corresponds to a real transaction. The mean home price, inflated by a few mansions, may be $500,000 and correspond to nothing at all.
The median isn’t a perfect statistic. It discards information about the tails, which can be important. But for the common purpose of describing a typical case, the median is superior to the mean in almost all situations involving skewed data, ordinal data, or data with undefined intervals. The responsible communicator reaches for the median first and only uses the mean when the distribution justifies it.
Why the Error Persists
If the problems with averaging are so well-documented, why do they persist? Several forces are at work. First, the mean is computationally simple and conceptually familiar. Everyone learned it in primary school. The median requires sorting the data, which is trivial with modern tools but still feels like an extra step. Second, the mean has mathematical properties that make it convenient for further statistical manipulation. Analysts who plan to run regressions or compute variances often start with means because the math works cleanly, even when the data does not.
Third, and most troubling, averages can be used to mislead. A politician who wants to claim rising prosperity can cite the mean income, knowing it’s pulled up by the wealthy. A school district can report a high average test score while hiding the bimodal distribution that reveals a stark achievement gap. The average, in these cases, is not a summary but a shield.
Media organizations share responsibility. The pressure to reduce complex stories to a single headline figure is immense. “Study Finds Average Screen Time Is Four Hours” is a simpler story than “Screen Time Varies Widely by Age, Income, and Device Type, With Distinct Patterns Across Groups.” The first headline fits a chyron; the second requires a paragraph. But the first headline is also false in the ways that matter most.
What the Careful Reader Should Do
When you encounter an average in the wild—in a news article, a political speech, a corporate report—pause. Ask what kind of data lies beneath it. Is it a mean or a median? Is the distribution symmetric or skewed? Are there subgroups that should be examined separately? Was the average computed from raw numbers or from percentages that should have been weighted?
If the report doesn’t provide this information, the average isn’t trustworthy. A single number without context is not a fact; it’s a rhetorical device. Demand the histogram, the box plot, the disaggregated table. If the source cannot or will not provide them, treat the average as an opinion, not a measurement.
Statistical literacy isn’t about mastering complex equations. It’s about recognizing when a number is being used to inform and when it’s being used to persuade. The average, misapplied, is one of the most common tools of persuasion masquerading as information.
FAQ
Why is the mean misleading for skewed data?
The mean is sensitive to extreme values. In a skewed distribution, a few very high or very low values pull the mean away from the bulk of the data. For example, in a neighborhood where most homes sell for around $300,000 but one sells for $3 million, the mean will be significantly higher than what a typical buyer experiences. The median remains close to $300,000 and better represents the central tendency.
Can ordinal data ever be averaged?
Strictly speaking, no. Ordinal data preserves order but not equal intervals. The difference between “Agree” and “Strongly Agree” is not necessarily the same as between “Neutral” and “Agree.” Computing a mean assumes equal spacing, which is rarely justified. The median is acceptable because it only uses order. For most purposes, reporting the full distribution of responses is more informative than any single summary statistic.
What is Simpson’s paradox and how does it relate to averaging?
Simpson’s paradox occurs when a trend that appears in several groups reverses when the groups are combined. This happens because the groups have different sizes and different baseline rates. The aggregate average weights the groups unequally and can produce a result that contradicts each group’s individual trend. It demonstrates that averaging across heterogeneous categories without accounting for composition can invert the truth.
How can I tell if an average in a news report is trustworthy?
Look for three things: whether the type of average (mean or median) is specified, whether the distribution is described or shown, and whether subgroups are discussed. A trustworthy report will mention if the data is skewed, will provide the median alongside the mean, and will break down results by relevant categories. If only a single average is given with no context, the number should be viewed with skepticism.