Numbers have a way of seducing us. They promise clarity, precision, a neat path through the mess. But there’s a quiet, stubborn error that slips into reports, dashboards, and news headlines more often than anyone wants to admit: averaging data that, by its very nature, should never be averaged. Jerome Leland has watched this mistake play out across finance, healthcare, public policy—and the fallout is rarely small.

The urge to average makes sense on the surface. A single figure feels clean. It slides into a headline, a bullet point, a ten-second sound bite without friction. But when the underlying data is built from ratios, percentages of mismatched wholes, or distributions that lean hard in one direction, the resulting “average” isn’t just misleading. It’s mathematically incoherent. It conjures a number that never actually existed in any real-world observation.

The Arithmetic Trap of Averages

At its heart, an average is a measure of central tendency. The arithmetic mean—sum everything up, divide by the count—works beautifully for additive quantities. Heights of ten people? Add them, divide by ten, and you get a meaningful central height. Same goes for weights, temperatures, or test scores where each observation stands on its own, measured on the same scale.

Trouble starts when the data points are themselves ratios, rates, or percentages drawn from different bases. Picture a small hospital that handles 10 surgeries and sees 1 complication—a 10% complication rate. A large hospital handles 1,000 surgeries and logs 50 complications, a 5% rate. Average 10% and 5% naively, and you get 7.5%. But the combined reality across both hospitals is 51 complications out of 1,010 surgeries—a rate of 5.05%. That 7.5% figure is a phantom. It describes no actual population of patients.

This error has a dramatic cousin called Simpson’s Paradox, where trends in subgroups flip when data gets aggregated improperly. But the underlying problem is simpler: averaging ratios without weighting them by their denominators produces a number detached from reality. Yet the practice is everywhere—in business metrics, educational statistics, even medical research summaries.

When Percentages Collide

Think about a marketing analyst comparing email open rates across five campaigns. Each campaign hits a different audience size. Campaign A reaches 100,000 recipients with a 20% open rate. Campaign B reaches 1,000 recipients with an 80% open rate. The simple average of 20% and 80% is 50%. But the actual combined open rate across both campaigns isn’t 50%—it’s yanked heavily toward the larger campaign. The true overall rate is (20,000 + 800) / (100,000 + 1,000) = 20.6%. Reporting 50% would be a catastrophic overstatement.

This isn’t academic. Jerome Leland has sat through investor pitch decks where customer retention rates were averaged across cohorts of wildly different sizes, producing a figure that suggested far healthier retention than actually existed. The startup’s largest cohort had a 40% retention rate, while a tiny pilot cohort showed 90%. The simple average of 65% masked the fact that the vast majority of customers were churning at a much higher rate.

Business professionals reviewing charts and data on a whiteboard
Data presentations often hide the denominator disparities that make simple averages misleading.

The Denominator Problem

At the bottom of this statistical misstep is a failure to respect the denominator. Every percentage, rate, or ratio is a fraction: a numerator divided by a denominator. The denominator represents the scale, the base, the population from which the numerator is drawn. When you average percentages without accounting for their denominators, you implicitly assume all denominators are equal. That assumption is rarely true and often wildly false.

Take crime rates across cities. City A has a population of 100,000 and reports 500 violent crimes, a rate of 5 per 1,000 residents. City B has a population of 10,000 and reports 100 violent crimes, a rate of 10 per 1,000. The simple average of the two rates is 7.5 per 1,000. But if you combine the cities, the actual rate is 600 crimes per 110,000 residents, or about 5.45 per 1,000. The unweighted average overstates the true rate by nearly 40%.

This distortion gets worse as the gap between denominators widens. When a small group with an extreme rate is averaged equally with a large group with a moderate rate, the small group punches above its weight. The result is a number that’s mathematically possible but empirically false—a synthetic statistic that represents no actual population.

Skewed Distributions and the Mean’s Deception

Another domain where averaging falls apart is highly skewed data. Income distributions, for example, are famously right-skewed: a small number of extremely high earners pull the arithmetic mean far above the median. Reporting the “average income” in such cases gives a distorted picture of what a typical person actually earns. The mean becomes a measure not of centrality but of total wealth divided by population, which is a different concept entirely.

In 2019, the U.S. Census Bureau reported a median household income of $68,703, while the mean household income was $98,088. That $30,000 gap isn’t a rounding error; it’s the gravitational pull of the top percentiles. When journalists or analysts casually refer to the “average American household income” without specifying which average, they often inadvertently misrepresent the economic reality for most households.

This problem extends to any metric with a long-tailed distribution: website page load times, hospital wait times, customer lifetime value, or the size of insurance claims. In each case, a few extreme values can inflate the arithmetic mean to a point where it no longer describes a typical experience. The median or geometric mean often provides a more honest central value, yet the arithmetic mean persists because it’s computationally simple and familiar.

When Averaging Destroys Information

Beyond the mathematical errors, there’s a deeper conceptual problem: averaging can destroy the very information that matters most. In public health, averaging infection rates across regions with vastly different population densities and healthcare access can hide dangerous hotspots. A national average might look reassuring while a specific city is experiencing an outbreak that requires immediate intervention.

In education, averaging test scores across schools with different demographics, funding levels, and class sizes can mask severe inequities. A district-wide average might suggest acceptable performance while individual schools are failing. The average becomes a tool of erasure, smoothing over the variation that policymakers most need to see.

In finance, averaging portfolio returns across time periods without considering volatility or sequence of returns can lead to disastrous planning. A retiree withdrawing funds during a market downturn faces a fundamentally different reality than the “average annual return” would suggest. The average hides the sequence-of-returns risk that can destroy a retirement plan.

Person analyzing financial charts and graphs on a laptop screen
Financial averages often conceal the volatility and sequence risk that determine real-world outcomes.

The Simpson’s Paradox Trap

Simpson’s Paradox is the most famous manifestation of the averaging error, and it deserves a closer look. The paradox occurs when a trend appears in several different groups of data but disappears or reverses when the groups are combined. The classic example comes from a 1973 study of gender bias in graduate admissions at the University of California, Berkeley. When looking at aggregate data, men appeared to be admitted at a higher rate than women. But when the data was broken down by department, women had equal or higher admission rates in most departments. The paradox arose because women tended to apply to more competitive departments with lower overall admission rates.

This is not a statistical curiosity; it’s a recurring trap in business analytics. A company might see that its overall customer satisfaction score is declining, yet every individual product line shows improving satisfaction. The paradox occurs because customers are migrating toward a product line with historically lower satisfaction scores. The aggregate trend points in the opposite direction of every subgroup trend. Acting on the aggregate number alone would lead to completely wrong conclusions.

The lesson isn’t that averages are useless, but that they must be applied with an understanding of the underlying data structure. When subgroups have different sizes and different base rates, the simple average is not just imprecise—it can be actively deceptive.

When Averaging Time-Series Data Misleads

Another common error is averaging data across time periods that have fundamentally different characteristics. Consider a retailer reporting “average monthly sales” over a year that includes both normal months and the holiday season. The average will be pulled upward by November and December, creating a figure that represents no actual month. A manager using this average to set monthly targets will find it too high for most of the year and too low for the peak season.

Similarly, averaging economic indicators across business cycle phases—recessions and expansions—produces a number that describes neither. The average GDP growth rate over a period that includes both a boom and a bust is a statistical artifact, not a meaningful benchmark. Yet such averages routinely appear in economic commentary as if they represented a stable underlying trend.

In medicine, averaging patient outcomes across different risk groups can obscure critical safety signals. A drug might show acceptable average efficacy, but when the data is stratified by age or comorbidity, it becomes clear that the drug is highly effective for some patients and dangerous for others. The average hides this heterogeneity, potentially leading to harmful prescribing practices.

Alternatives to the Naive Average

When faced with data that should not be simply averaged, several alternatives exist. The choice depends on what question is actually being asked.

Weighted averages correct the denominator problem by giving each data point influence proportional to its size. For the hospital complication rates, weighting by the number of surgeries at each hospital produces the correct aggregate rate. This is the most direct fix for the ratio-averaging error and should be the default approach whenever combining rates or percentages from groups of different sizes.

Medians resist the pull of extreme values and better represent the “typical” observation in skewed distributions. For income data, home prices, or any metric with a long tail, the median is almost always more informative than the mean for describing what a typical person or case experiences.

Geometric means are appropriate for averaging growth rates, returns, or other multiplicative processes. If an investment grows 50% one year and falls 50% the next, the arithmetic mean return is 0%, but the geometric mean correctly shows a net loss. The arithmetic mean of percentage changes over time is almost always misleading.

Stratified reporting preserves the variation that simple averaging destroys. Instead of reporting a single average, report the averages for each meaningful subgroup alongside the overall weighted average. This approach maintains transparency and allows readers to see both the forest and the trees.

Multiple charts and graphs displayed on a wall showing different data segments
Stratified reporting preserves the variation that simple averages erase.

Real-World Consequences

The consequences of averaging data that should not be averaged extend far beyond academic debates. In 2016, a major pharmaceutical company presented aggregated clinical trial results that showed a modest benefit for a new drug. Only when independent researchers reanalyzed the data by patient subgroups did it become clear that the drug was actively harmful to patients over 65 while providing significant benefit to younger patients. The simple average had masked a critical safety signal.

In public policy, averaging unemployment rates across regions with vastly different economic structures can lead to misallocation of resources. A national unemployment rate of 5% might sound healthy, but if it consists of 2% unemployment in urban areas and 15% in rural communities, the average hides a crisis. Programs designed based on the national average will fail to address the real problem.

In education, the No Child Left Behind Act required schools to report average test scores, which led to well-documented instances of schools focusing resources on students near the proficiency threshold while neglecting both high-achieving and severely struggling students. The average became the target, and the target distorted the system.

How to Spot the Error

Readers and analysts can protect themselves by asking a few diagnostic questions whenever they encounter an average:

First, what is being averaged? If the answer is a percentage, rate, or ratio, immediately check whether the denominators are equal. If they are not, the average is almost certainly wrong unless it has been explicitly weighted.

Second, what is the distribution? If the data is skewed, the mean will be pulled away from the typical value. Ask for the median or a full distribution plot. A single number cannot capture a skewed reality.

Third, what variation is being hidden? Even a correctly calculated weighted average can conceal important differences between subgroups. Always ask for the disaggregated data. If the source cannot or will not provide it, treat the average with skepticism.

Fourth, is the average being used to make a decision that depends on the full distribution? In risk management, inventory planning, or capacity sizing, the average is often insufficient. The extremes—the tails of the distribution—frequently determine success or failure. A bridge designed to withstand the average flood will collapse regularly.

FAQ

Why can’t I just average percentages from different groups?

Percentages are ratios with denominators that often differ in size. Averaging them without weighting by those denominators gives equal influence to groups of unequal size, producing a number that does not correspond to the actual combined rate. For example, averaging a 90% rate from a group of 10 and a 40% rate from a group of 1,000 yields 65%, but the true combined rate is about 40.5%. The simple average is mathematically incoherent because it ignores the scale of each group.

When is it acceptable to use a simple average?

A simple arithmetic mean is appropriate when the data points are additive quantities measured on the same scale and each observation carries equal weight. Examples include averaging the heights of individuals in a room, the temperatures recorded at different weather stations on the same day, or the scores of students who all took the same test. The key condition is that each data point represents a measurement of the same kind with a consistent underlying unit.

What should I use instead of a simple average for skewed data?

For skewed distributions, the median is usually a better measure of central tendency because it is not pulled by extreme values. For growth rates or returns over time, the geometric mean is more appropriate because it accounts for compounding effects. For data with unequal group sizes, a weighted average that uses the group sizes as weights will give the correct aggregate figure. In many cases, reporting the full distribution or providing stratified averages alongside the overall figure is the most honest approach.

How does averaging cause bad business decisions?

Averaging can lead to bad decisions when it hides variation that matters for planning. If a retailer uses average monthly sales to set inventory levels, they will be overstocked in slow months and understocked in peak months. If an investor uses average annual returns to plan retirement withdrawals without considering the sequence of those returns, they may run out of money during an early downturn. The average provides a single number, but real-world outcomes depend on the entire pattern of variation.

The Hidden Danger of Averaging Data That Should Never Be Averaged