How to Track Pandemic Data Without the Sensationalism

Why Most Pandemic Dashboards Overwhelm Instead of Informing

When COVID-19 first swept the globe, a strange new pastime took hold: watching the numbers. Dashboards, graphs, and real-time counters became daily obsessions. But for plenty of people, the flood of data—case counts, hospitalizations, deaths, positivity rates, Rt values—created more fog than clarity. The issue wasn’t a shortage of information. It was a shortage of context. Raw figures, stripped of explanation, can easily mislead. And sensationalist headlines? They almost always grab the most frightening stat and run with it. Tracking pandemic data without the noise means learning to read the signals buried in the static.

Jerome Leland has spent years sifting through public health data for community news outlets, and the first lesson he learned is straightforward: numbers alone rarely tell the whole story. A sudden jump in cases might just mean more tests were processed, not that infections exploded. A dip in hospitalizations could be a reporting delay, not a sign the wave is over. The trick is to build a mental framework that filters out panic and zeroes in on what the data is actually saying.

Person analyzing data charts on a laptop screen

Start With the Right Sources

All data is not created equal. Official public health agencies—the CDC, the WHO, your state or county health department—are still the best starting points. They publish raw numbers, explain their methods, and issue regular situation reports. Avoid dashboards that just flash a big red number with no explanation. Those are built for emotional punch, not understanding.

What you want is metadata: how cases are defined, what testing criteria are used, whether probable cases are lumped in with confirmed ones. During the Omicron wave, for instance, some jurisdictions stopped emphasizing case counts and focused on hospitalizations instead. Why? Because so many breakthrough infections in vaccinated people were mild, the case numbers painted a scarier picture than reality. Without that context, a viewer might see a sudden drop in reported cases and assume the wave had crashed—when really, the surveillance system had just shifted its lens.

Know Which Metrics Matter—and When

Not every metric is equally useful at every stage of a surge. Early on, case counts and test positivity can give you a heads-up that something is brewing. Later, hospital admissions and ICU occupancy tell you how bad it really is. Deaths are the ultimate lagging indicator, often peaking two to four weeks after cases. If you track all of them without understanding their timing, you’ll get whiplash: one day the news looks catastrophic, the next it looks like a miracle, when the underlying trend hasn’t budged.

Here’s a cheat sheet for the most common metrics:

  • Case counts: Good for spotting early trends, but they swing wildly based on testing access and reporting delays. Treat them as a rough compass, not a precise map.
  • Test positivity rate: The percentage of tests coming back positive. Above 5% suggests you’re missing cases; below 2% means testing is probably adequate. A simple but powerful gauge.
  • Hospital admissions: A steadier measure of severe illness, less influenced by who decides to get tested. Pay attention to age shifts—if younger people start filling beds, the variant’s behavior may be changing.
  • ICU occupancy: The most serious cases end up here. When ICU beds are scarce, the healthcare system is strained, no matter what the case counts say.
  • Deaths: The final, tragic confirmation of what happened weeks ago. Essential for understanding overall impact, but useless for real-time decisions.
  • Wastewater surveillance: A newer tool that detects viral RNA in sewage. It doesn’t care about testing access or human behavior—it’s a leading indicator of community spread that’s gaining traction.

Close-up of a person reading a chart on a tablet

Spotting Misleading Visuals

Data visualization can illuminate—or it can trick you. One common sleight of hand is truncating the y-axis on a graph. If a bar chart starts at 900 instead of zero, a modest 5% bump can look like a terrifying spike. Always glance at the axis scales before reacting. Another trap: comparing raw numbers between places with wildly different populations. Fifty cases in a town of 10,000 is a serious outbreak; 200 cases in a city of a million is far less alarming. Per capita rates—cases per 100,000 residents—are your friend for fair comparisons.

Color choices sneak into your brain, too. Heat maps that splash intense reds over relatively low case rates can trigger unnecessary dread. Meanwhile, green-shaded maps might whisper “all clear” when transmission is still humming along. Look for visuals with clear, graduated scales and the ability to hover or click for the actual numbers. If a chart makes you feel something before you understand something, take a breath and examine the design choices.

Building a Personal Tracking Routine

Jerome Leland suggests a weekly check-in for most people, not a daily one. Daily numbers are noisy—they bounce around because of reporting lags, weekend slowdowns, and occasional data dumps. A seven-day rolling average smooths out those wrinkles and shows the real direction. Pick two or three metrics that matter to you—maybe hospital admissions and wastewater levels—and stick with them. Write the numbers down each week. The simple act of recording forces you to engage with the data instead of passively absorbing headlines.

Here’s a straightforward template:

  1. Find a reliable source for your area: your state health department, county dashboard, or the CDC’s COVID Data Tracker.
  2. Choose two or three metrics that fit your personal risk picture. If you’re immunocompromised, wastewater levels and test positivity might be your go-tos. If you’re watching healthcare capacity, track hospital admissions and ICU occupancy.
  3. Record the numbers once a week, on the same day. Jot down the date and any methodological tweaks the source announces.
  4. Watch the trend over four to six weeks. A single week’s change is noise; a month’s direction is signal.

This habit builds data literacy and a kind of emotional steadiness. When a scary headline screams at you, you can glance at your own records and judge whether the story matches the trend you’ve been watching.

Putting Variants and Vaccines in Context

New variants generate headlines that often sprint ahead of the science. When a variant gets tagged “of concern,” it means preliminary studies suggest it’s more transmissible, more severe, or better at dodging immunity. But “preliminary” is the operative word. Plenty of variants that set off alarm bells fizzled out because they couldn’t outcompete the dominant strain. Track variant proportions through genomic surveillance dashboards—the CDC and GISAID maintain good ones—but remember: a variant’s rise in sequencing data might just reflect more sampling, not necessarily more transmission.

Vaccine effectiveness data also demands a careful read. Studies often report effectiveness against infection, symptomatic disease, hospitalization, and death separately. A headline blaring “vaccine effectiveness drops to 40%” is usually talking about infection alone. Protection against severe outcomes typically stays much higher. Breakthrough cases among vaccinated people are normal and don’t mean the vaccines have failed. The metric that matters is the rate of severe outcomes in vaccinated versus unvaccinated groups, adjusted for age and underlying conditions.

Person wearing a mask while looking at a smartphone

Recognizing Sensationalism in Headlines

Media outlets are under pressure to grab clicks, and fear is a reliable seller. Sensationalist pandemic coverage tends to follow a few predictable patterns:

  • Absolute numbers without denominators: “10,000 new cases!” with no mention of the population or how many tests were run.
  • Cherry-picked timeframes: Comparing a terrible day to a quiet one instead of showing the seven-day average.
  • Emotionally loaded language: Words like “skyrocket,” “plummet,” “terrifying,” or “miracle” are editorializing, not reporting.
  • Single-source stories: Articles that hang their entire argument on one preprint or one expert’s opinion, ignoring the broader scientific consensus.
  • Missing denominators in risk comparisons: “Triple the risk” sounds huge, but if the absolute risk is still tiny, it’s a nothingburger.

When you bump into a dramatic claim, pause and ask: What’s the absolute risk? What’s the baseline? What’s the trend over time? Who is actually affected? A level-headed reader treats each headline as a hypothesis to check, not a conclusion to swallow.

FAQ: Common Questions About Pandemic Data

Why do case numbers sometimes drop suddenly?

Sudden drops in reported cases usually reflect changes in reporting practices, not true declines in transmission. Common culprits: reduced testing availability, tweaks to case definitions, backlogs getting cleared, or holidays delaying data entry. Always look for methodological notes from the reporting agency before you interpret a sharp change as a real trend.

What’s the difference between case fatality rate and infection fatality rate?

The case fatality rate (CFR) is confirmed deaths divided by confirmed cases. It overestimates lethality because many mild or asymptomatic infections never get confirmed. The infection fatality rate (IFR) estimates deaths among all infections, including undiagnosed ones. IFR is usually much lower than CFR but harder to calculate—it requires seroprevalence studies or modeling to estimate total infections. When you see a “death rate” quoted, check which denominator is being used.

How can I tell if a data source is trustworthy?

Trustworthy sources are transparent about their methods. They explain how data is collected, what definitions they use, and what limitations exist. They update regularly and correct errors publicly. Official public health agencies, academic institutions, and established research organizations generally meet these standards. Be wary of sources that present data without methodology, use emotionally manipulative design, or have a clear advocacy agenda that might slant their presentation.

Putting It All Together

Tracking pandemic data without the sensationalism is a skill that sharpens with practice. It’s about choosing credible sources, understanding what each metric actually measures, spotting misleading visuals, and sticking to a consistent personal routine. The payoff is a clearer view of what’s happening in your community—and the confidence to make decisions based on evidence, not fear. In a media environment that often prizes alarm over accuracy, that clarity is worth the effort.

What Happened When We Fed State Unemployment Data to an AI Script Generator

Every data journalist knows the moment. You have a clean dataset, a clear chart, and a finding worth sharing. Then comes the hard part: building a narrative that respects the numbers without turning them into a bedtime story. The temptation to reach for a tool like an AI script generator—such as an AI script generator that fits the project—is real, and growing. These tools ingest a spreadsheet or a statistical summary and output a structured script, complete with scene headings, voiceover cues, and dramatic arcs. They promise speed. They promise structure. What they rarely promise, and what journalists must demand, is fidelity to the data.

This article is not a blanket condemnation of automation. It is a method for evaluating what automation produces. The same rigor we apply to chart design—checking axes, questioning baselines, disclosing uncertainty—must extend to the narrative wrapper. Because a script is not neutral. It selects, sequences, and emphasizes. When that selection is done by a model trained on generic story patterns rather than on the specific dataset in front of you, the result can flatten uncertainty, invent causality, and favor the most dramatic available arc. The question is not whether to use these tools. The question is how to read their output with the skepticism of a data editor.

The Case Study: State-Level Unemployment, Month by Month

To make this concrete, I worked with a real dataset: monthly state-level unemployment rates from the Bureau of Labor Statistics, covering January 2023 through March 2026. The data includes all 50 states plus the District of Columbia. It is seasonally adjusted. It is publicly available. It is exactly the kind of dataset a newsroom might use to build a regional economic story or a comparative feature on labor market recovery.

I then did two things. First, I wrote a short narrative script by hand—the kind a data journalist might draft for a three-minute video segment or an annotated scrolling feature. Second, I fed the same dataset and a simple prompt into a commercially available AI script generator and examined what came back. The specific tool used was the Unsloppy AI script generator, a web-based service that accepts structured data and a topic prompt, then returns a formatted screenplay-style narrative. Its public documentation is limited, and the terms of use do not restrict naming it in an editorial evaluation. The comparison is instructive not because the human version is perfect, but because the differences reveal systematic weaknesses in automated narrative construction.

The Human-Crafted Script

Here is the core of the hand-written script, condensed for print:

SCENE 1: WIDE SHOT — U.S. MAP, JANUARY 2023
VOICEOVER: In January 2023, the national unemployment rate sat at 3.4 percent. But national averages hide as much as they reveal. Nevada was at 5.2. Maryland at 2.1. That spread—3.1 percentage points—is the story we’re going to track.

SCENE 2: LINE CHART — FOUR STATES, 2023–2026
VOICEOVER: We selected four states to follow: Nevada, Maryland, Michigan, and Texas. Not because they’re extreme, but because they represent different labor market structures. Nevada is hospitality-heavy. Maryland has a high federal employment share. Michigan is manufacturing-exposed. Texas is large and energy-adjacent. Watch what happens.

SCENE 3: CLOSE-UP — NEVADA LINE DROPS, MICHIGAN LINE FLAT
VOICEOVER: Nevada’s rate fell from 5.2 to 4.1 over three years. That’s a genuine improvement. Michigan’s rate barely moved—4.1 to 4.0. But here’s what the line doesn’t show: Michigan’s labor force shrank by 1.2 percent over the same period. A stable unemployment rate with a shrinking labor force is not stability. It’s a warning.

SCENE 4: SPLIT SCREEN — TEXAS AND MARYLAND
VOICEOVER: Texas added 900,000 jobs. Its unemployment rate rose slightly—from 3.9 to 4.2—because the labor force grew even faster. That’s a healthy dynamic. Maryland’s rate stayed low, but its labor force participation rate declined. Two low unemployment numbers. Two very different stories.

SCENE 5: FULL MAP — MARCH 2026, WITH ARROWS
VOICEOVER: The national rate in March 2026 was 3.8 percent. The state spread narrowed to 2.4 points. But narrowing can mean convergence or it can mean the high-rate states improved while the low-rate states deteriorated. In this case, it was mostly the former. But you can’t know that from the spread alone.

SCENE 6: TITLE CARD — “WHAT WE DON’T KNOW”
VOICEOVER: These numbers are seasonally adjusted. That means the BLS has removed predictable calendar effects. But seasonal adjustment is a model. It revises. The January 2026 figure you see today may shift by 0.2 points when the annual revision arrives next year. Also, state-level samples are smaller than national samples. The error range on Nevada’s 4.1 is wider than on the U.S. 3.8. We’re showing point estimates. The data is trying to tell you there’s fog around every number. Listen for it.

Notice what this script does. It names its selection criteria. It distinguishes between a falling unemployment rate that signals recovery and a stable rate that masks labor force contraction. It explains seasonal adjustment and revision risk. It ends on uncertainty, not on a tidy conclusion. The narrative arc is present—states change, patterns shift—but the arc is subordinate to the data’s internal logic.

The AI-Generated Script

The automated script, by contrast, opened with a different move:

SCENE 1: DRAMATIC MUSIC — DARKENED MAP, FLASHING RED STATES
VOICEOVER: In the wake of economic turmoil, America’s job market faced a crisis. Some states burned while others thrived. This is the story of a divided nation, told through the numbers that matter most.

The generator had detected variance in the dataset and interpreted it as crisis. It had no access to the fact that a 3.1-percentage-point spread in unemployment during a period of 3.4 percent national unemployment is historically narrow, not wide. It saw difference and reached for the most dramatic available frame: division, turmoil, burning versus thriving.

From there, the automated script selected the five states with the highest and lowest unemployment rates in January 2023 and tracked only those. It ignored the middle. It constructed a narrative of “recovery leaders” and “lagging states” without once mentioning labor force participation, industry mix, or sampling error. It introduced causal language: “Nevada’s tourism-dependent economy dragged it down,” “Texas boomed thanks to energy expansion.” The dataset contains no industry-level variables. The model inferred causality from correlation with geographic stereotypes.

The script closed with a triumphant scene: “By March 2026, the gap had closed. America had healed.” The narrowing spread was presented as resolution. The possibility that low-rate states had simply drifted upward—a less satisfying story—was absent. The uncertainty scene from the human script had no equivalent. There was no mention of seasonal adjustment, revision, or sample size.

Where Automation Fails: A Pattern Recognition Problem

What happened here is not unique to this dataset or this tool. It is a structural tendency of automated narrative generators trained on large corpora of human-written stories. These models learn that a story has a beginning, a middle, and an end. They learn that conflict and resolution are preferred. They learn that extreme values are more narratively useful than central tendencies. They do not learn—because the training data rarely teaches—that the most honest data story sometimes concludes with “we cannot yet say.”

The Authors Guild, in its AI Best Practices for Authors, warns that “AI outputs are generic mashups of pre-existing works ingested during training” and that “when you claim authorship in a work, it means” you are responsible for the thinking behind it. That warning applies with special force to data journalism. A script that flattens uncertainty and invents causality is not just aesthetically generic. It is epistemically damaging. It trains audiences to expect false clarity.

Professional screenwriting, as StudioBinder’s guide on how to write a movie script explains, relies on deliberate structural choices: scene headings that convey geography and time, subheadings that signal shifts without breaking flow, formatting that ensures clarity. These conventions exist to serve the story, not to replace it. An AI script generator can reproduce the format—it can insert INT. and EXT., it can label scenes—but it cannot make the deliberate choices that give those labels meaning. It cannot decide that a scene heading should read “CLOSE-UP — NEVADA LINE DROPS, MICHIGAN LINE FLAT” because that specific juxtaposition carries analytical weight. It will instead generate generic labels: “SCENE 3: THE STRUGGLING STATE” or “SCENE 4: THE SUCCESS STORY.” The format is present. The thinking is absent.

A Checklist for Evaluating Any Data-Driven Script

The solution is not to reject automation. It is to read automated output with the same structured skepticism we bring to a chart with a truncated y-axis or an unlabeled color scale. Below is a checklist designed for newsrooms, educators, and anyone who encounters a script or storyboard that claims to be “generated from data.” Run through these questions before you publish, air, or cite.

  1. Does the script name its selection criteria? If it highlights specific cases (states, sectors, demographic groups), can you explain why those and not others? A script that picks extremes without disclosing the selection rule is building drama on an unstated foundation.
  2. Does the script distinguish between correlation and causation? Look for phrases like “led to,” “drove,” “caused,” “triggered.” If the underlying dataset contains only observational variables, causal language is an inference the script has added. Demand to see the evidence for that inference, or strike the language.
  3. Does the script acknowledge what the dataset cannot show? Every dataset has limits: sample size, measurement error, revision schedules, unmeasured confounders. A script that never says “we don’t know” or “this estimate will revise” is overselling its certainty.
  4. Does the narrative arc match the data’s internal shape, or is it imposed? Not every dataset has a climax. Some trends are flat. Some cycles are incomplete. If the script forces a three-act structure onto data that is essentially a plateau with noise, the structure is lying.
  5. Are the visual cues honest? If the script calls for “flashing red states” or “dramatic music,” ask what in the data justifies that emotional register. A 0.3-percentage-point change in a rate with a 0.2-point error margin does not justify an alarm.
  6. Does the script define its terms? “Unemployment rate” means something specific. “Labor force participation rate” means something else. If the script uses terms without defining them, it assumes a level of literacy the audience may not have—and may exploit that gap for dramatic effect.
  7. Does the script disclose its source and vintage? “BLS data” is not enough. Which survey? Which seasonal adjustment? Which release month? Data vintages change. A script built on preliminary January data may be obsolete by March. The script should tell the audience when the data was pulled and that it may revise.
  8. Does the script include a scene for uncertainty? If not, add one. Even a 15-second title card that says “These numbers have error margins. The state-level estimates are less precise than the national figure. Revisions will come.” That scene is not a buzzkill. It is the difference between journalism and theater.

Why the Narrative Wrapper Matters as Much as the Chart

For years, data journalism has focused—rightly—on visual honesty. We have debated truncated y-axes, misleading color scales, and the sins of the dual-axis chart. But the narrative that surrounds the chart is equally capable of distortion. A perfectly honest line chart, placed inside a script that invents causality and hides uncertainty, becomes a prop in a fiction. The audience remembers the story, not the axis labels.

This is why the checklist above is not optional. It is the narrative equivalent of requiring labeled axes, cited sources, and visible error bars. A newsroom that would never publish a chart without a source line should not publish a script without a methodological disclosure. A journalism school that teaches students to critique visualizations should also teach them to critique the storyboards that those visualizations inhabit.

The AI script generator is not the enemy. It is a tool that accelerates a process—drafting narrative structure—that journalists have always done. The risk is not that the tool exists. The risk is that we accept its output without the same scrutiny we apply to every other stage of data reporting. We would not publish a chart generated by an AI that had never seen the underlying data. We should not publish a script generated by an AI that has seen the data but not understood it.

Toward a Shared Standard for Data Narratives

What would a shared standard look like? Transparency is the starting point. Any published data script—whether human-written, AI-assisted, or fully automated—should carry a brief disclosure: how the narrative was constructed, what tool or process was used, what human editorial review was applied. This is not radical. Film scripts carry credits. Investigative reports carry methodology notes. Data narratives should carry the same.

Training is the next step. Newsrooms that adopt AI script generators should train their staff to use the checklist above, just as they train them to spot a misleading y-axis. The skill of reading a script for epistemic honesty is teachable. It requires no advanced statistics. It requires only the habit of asking: “How does this sentence know what it claims to know?”

And then there is the cultural shift. The most honest data story is sometimes the one that says: “Here is what we see. Here is what we cannot see. Here is what would change our interpretation.” That story does not end with a triumphant chord. It ends with an invitation to keep looking. That is not a failure of narrative. It is the defining virtue of data journalism. The question is whether newsrooms are ready to treat that virtue as a requirement, not an aspiration.

The next time you encounter a script that claims to be “generated from data,” run the checklist. If it fails more than two items, it is not a data story. It is a story that happens to have some numbers in it. The difference matters. In a media environment saturated with automated content, the ability to tell the difference is a skill worth building—one that the checklist above is designed to support.

The Problem With Averaging Data That Should Not Be Averaged

We love a simple number. One figure that can wrap up a spreadsheet, a trend, a person, a policy. The average is the tool we grab first. It feels fair, sanding down the extremes into a comfortable middle. But that comfort is often a statistical mirage, and when you force it onto data that resists being averaged, you don’t get clarity. You get distortion. The trouble isn’t the arithmetic. It’s the assumption that every dataset has a meaningful center.

Jerome Leland has spent decades watching how numbers get used and misused in public conversation. The mistake isn’t always malicious. Sometimes it comes from rushing, sometimes from a wish to make complicated things digestible for an audience with a short attention span. But the consequences are real. When we average what should not be averaged, we erase the exact information we need to make sound decisions.

The Arithmetic Mean: A Tool With Limits

The arithmetic mean—sum the values, divide by the count—gets taught early and used everywhere. It’s the perfect summary for data that’s symmetric and unimodal, like the heights of adult men in a specific population, or the manufacturing tolerances of a well-calibrated machine. In those cases, the average sits near the center of a bell-shaped distribution, and most individual values cluster around it. The average is representative.

But the world is full of data that isn’t symmetric. Income distributions, for example, are heavily right-skewed. A small number of extremely high earners yank the arithmetic mean upward, far away from what a typical person actually earns. Reporting the average income without disclosing the median—the point where half earn more and half earn less—is a classic, and often deliberate, statistical misdirection. The average becomes a tool for hiding inequality rather than revealing a central tendency.

This isn’t a subtle point, yet it gets routinely ignored in headlines and political speeches. The problem isn’t that the average is mathematically wrong; it’s that it answers a question nobody asked. When people hear “average,” they imagine the typical case. For skewed data, the arithmetic mean describes no actual case at all.

Ordinal Data: When Numbers Are Just Labels

An even more basic error happens when we average data that isn’t truly numerical. Ordinal data—rankings, ratings, categories assigned numbers for convenience—carries order but not magnitude. The difference between a 1-star and a 2-star rating is not the same as the difference between a 4-star and a 5-star rating. The numbers are placeholders, not measurements.

Consider a customer satisfaction survey where respondents choose from “Very Dissatisfied” (coded as 1), “Dissatisfied” (2), “Neutral” (3), “Satisfied” (4), and “Very Satisfied” (5). Calculating a mean score of 3.7 is common practice, but it’s statistically indefensible. The result implies a continuous scale with equal intervals, which doesn’t exist. A score of 3.7 doesn’t mean the average respondent is slightly more than “Neutral.” It’s a numerical ghost, a fiction created by treating labels as measurements.

The proper summary for ordinal data is the median or the mode, and the distribution should be reported in full. A product with fifty 5-star ratings and fifty 1-star ratings has an average of 3, but that average conceals a deeply polarized reception. The mean tells a story of mediocrity; the distribution tells a story of controversy. Which story is more useful? Almost always, the latter.

Bimodal and Multimodal Distributions: The Average That Describes Nobody

Some datasets have two or more distinct peaks. The classic example is human height when sexes are combined. Adult men and women have overlapping but distinctly different height distributions. The overall average height for a mixed-gender population falls somewhere between the male and female averages, and it describes almost no one. It’s a statistical artifact, not a human being.

This problem appears in more subtle forms across media reporting. Consider a study on screen time that lumps together children under five, school-age children, and teenagers. The average screen time might be reported as four hours per day, but that number masks enormous variation. A toddler watching educational videos for thirty minutes gets averaged with a teenager spending eight hours on social media. The resulting figure is useless for policy, useless for parenting advice, and actively misleading for public understanding.

When data is multimodal, the responsible approach is to report the modes separately. Describe the clusters. Acknowledge that the population isn’t homogeneous. The average, in these cases, is a lie of aggregation.

A diverse group of people in a meeting, illustrating the danger of averaging distinct subgroups into a single meaningless number

Time-Series Data and the Ecological Fallacy

Averaging over time introduces its own species of error. A stock price that fluctuates wildly between $10 and $100 may have an average of $55, but that number tells you nothing about the risk an investor faced. The sequence matters. The volatility matters. The average smooths away the very information that drives decisions.

In climate science, this error has profound implications. Global average temperature is a carefully constructed metric, but it’s often misinterpreted. A rise of 1.5°C in the global average doesn’t mean every place gets 1.5°C warmer. Some regions warm much more; others may temporarily cool. The average conceals regional extremes, shifts in precipitation, and changes in seasonality. When a policymaker says “we can adapt to 1.5 degrees,” they’re adapting to a number, not to the lived reality of heatwaves, floods, and crop failures that the average hides.

The ecological fallacy—assuming that group-level statistics apply to individuals within the group—is a constant companion to improper averaging. A county with an average income of $80,000 may contain neighborhoods of deep poverty and enclaves of extreme wealth. The average is true for the county as a unit, but false for nearly every person in it.

Proportions and Percentages: The Unweighted Mean Trap

Another common blunder is taking the average of percentages without weighting by the base. If a small hospital has a 100% patient satisfaction rate from 10 patients, and a large hospital has an 80% rate from 1,000 patients, the unweighted average is 90%. This suggests a level of satisfaction that isn’t remotely representative of the patient experience. The vast majority of patients were 80% satisfied, but the small hospital’s perfect score—based on a tiny sample—drags the average upward.

This error appears routinely in school district comparisons, product reviews, and medical success rates. The fix is simple: always weight by the denominator. But the simple fix is often ignored because the unweighted average is easier to compute and, sometimes, because it produces a more favorable number for the party doing the reporting.

When Averages Become Dangerous: Policy and Medicine

The consequences of averaging inappropriate data move from misleading to dangerous when they inform policy or medical decisions. Drug trials that report average efficacy may conceal the fact that a medication works well for some patients, has no effect on others, and harms a small subset. The average benefit looks positive, but the individual experience is a lottery. Personalized medicine exists precisely because averages fail to capture this heterogeneity.

In public health, averaging infection rates across a large geographic area can hide hotspots that need immediate intervention. During the early months of the COVID-19 pandemic, national averages obscured the catastrophic situation in specific cities and nursing homes. Decision-makers who relied on the average delayed action; those who looked at the distribution acted faster and saved lives.

Economic policy suffers from the same flaw. Average wage growth can be positive while the median wage stagnates or falls, if gains are concentrated at the top. Reporting the average without the median gives citizens and lawmakers a false sense of progress. The average becomes a tool for maintaining the status quo, not for measuring the well-being of the typical person.

A stethoscope resting on a financial report, symbolizing the intersection of health and economic data where averaging can mislead

When the Average Is Not Even a Number

Some data cannot be averaged because the result is not just misleading—it’s meaningless. Consider categorical data like marital status, blood type, or political party affiliation. Assigning numbers to these categories and computing a mean produces absurdity. The average blood type is not a useful concept. The average political party is not a moderate; it’s a statistical ghost.

Yet this kind of averaging creeps into survey reporting. Researchers assign numerical codes to responses for convenience in data entry, then forget that the codes are not quantities. A mean of 2.4 for religious affiliation (where 1=Christian, 2=Muslim, 3=Jewish, etc.) is not a person who is slightly more Muslim than Christian. It’s a category error, a confusion of measurement scales that renders the result uninterpretable.

Nominal data—categories without order—must never be averaged. Ordinal data—rankings without equal intervals—should not be averaged, though the median may sometimes be acceptable. Only interval and ratio data, where differences are meaningful and a true zero exists, can be properly averaged. Even then, the distribution must be checked for skew and outliers.

The Simpson’s Paradox: When Averages Reverse Truth

Simpson’s paradox is one of the most startling demonstrations of how averaging can invert reality. It occurs when a trend appears in several different groups of data but disappears or reverses when these groups are combined. The classic example involves university admissions: within each department, women may have a higher acceptance rate than men, but overall, men have a higher acceptance rate because women apply in larger numbers to departments with lower acceptance rates. The aggregate average tells the opposite story of the disaggregated data.

This paradox appears in medical studies, social science research, and business analytics. A marketing campaign may show improved conversion rates in every segment, yet the overall conversion rate drops because the mix of segments shifted toward lower-converting groups. Averaging across segments without accounting for composition changes produces a false narrative of decline.

The only defense against Simpson’s paradox is to examine data at the appropriate level of granularity. Averages computed across heterogeneous groups are not summaries; they are distortions.

Practical Guidance: When to Avoid the Average

For anyone working with data—journalists, analysts, policymakers, or curious citizens—a few simple checks can prevent the misuse of averages. Before computing a mean, ask these questions:

  • Is the distribution symmetric? Plot a histogram. If the data is skewed, the mean will be pulled toward the tail and will not represent the typical value. Use the median or report the full distribution.
  • Are there multiple peaks? If the data clusters into distinct groups, the mean falls in an unpopulated valley. Describe the clusters separately.
  • Is the data ordinal or nominal? If the numbers are just labels, do not compute a mean. Use the mode or report category frequencies.
  • Are the groups of different sizes? If you are averaging percentages or rates from subgroups, weight by the subgroup size. An unweighted mean of percentages is almost always wrong.
  • Does the average answer the question you are actually asking? If you want to know what is typical, the mean may not tell you. If you want to know the total impact, the mean multiplied by the count may be appropriate, but only if the data is ratio-scale and the distribution is understood.

These checks take seconds. They prevent errors that can last for years.

A person analyzing data on a whiteboard with charts and graphs, emphasizing the need for careful statistical reasoning

The Median: A Better Default

For most real-world data that reaches the public, the median is a more honest summary than the mean. The median is resistant to outliers, works for ordinal data, and always represents an actual data point when the sample size is odd. It doesn’t create fictional values. When a journalist reports that the median home price in a city is $350,000, that number corresponds to a real transaction. The mean home price, inflated by a few mansions, may be $500,000 and correspond to nothing at all.

The median isn’t a perfect statistic. It discards information about the tails, which can be important. But for the common purpose of describing a typical case, the median is superior to the mean in almost all situations involving skewed data, ordinal data, or data with undefined intervals. The responsible communicator reaches for the median first and only uses the mean when the distribution justifies it.

Why the Error Persists

If the problems with averaging are so well-documented, why do they persist? Several forces are at work. First, the mean is computationally simple and conceptually familiar. Everyone learned it in primary school. The median requires sorting the data, which is trivial with modern tools but still feels like an extra step. Second, the mean has mathematical properties that make it convenient for further statistical manipulation. Analysts who plan to run regressions or compute variances often start with means because the math works cleanly, even when the data does not.

Third, and most troubling, averages can be used to mislead. A politician who wants to claim rising prosperity can cite the mean income, knowing it’s pulled up by the wealthy. A school district can report a high average test score while hiding the bimodal distribution that reveals a stark achievement gap. The average, in these cases, is not a summary but a shield.

Media organizations share responsibility. The pressure to reduce complex stories to a single headline figure is immense. “Study Finds Average Screen Time Is Four Hours” is a simpler story than “Screen Time Varies Widely by Age, Income, and Device Type, With Distinct Patterns Across Groups.” The first headline fits a chyron; the second requires a paragraph. But the first headline is also false in the ways that matter most.

What the Careful Reader Should Do

When you encounter an average in the wild—in a news article, a political speech, a corporate report—pause. Ask what kind of data lies beneath it. Is it a mean or a median? Is the distribution symmetric or skewed? Are there subgroups that should be examined separately? Was the average computed from raw numbers or from percentages that should have been weighted?

If the report doesn’t provide this information, the average isn’t trustworthy. A single number without context is not a fact; it’s a rhetorical device. Demand the histogram, the box plot, the disaggregated table. If the source cannot or will not provide them, treat the average as an opinion, not a measurement.

Statistical literacy isn’t about mastering complex equations. It’s about recognizing when a number is being used to inform and when it’s being used to persuade. The average, misapplied, is one of the most common tools of persuasion masquerading as information.

FAQ

Why is the mean misleading for skewed data?

The mean is sensitive to extreme values. In a skewed distribution, a few very high or very low values pull the mean away from the bulk of the data. For example, in a neighborhood where most homes sell for around $300,000 but one sells for $3 million, the mean will be significantly higher than what a typical buyer experiences. The median remains close to $300,000 and better represents the central tendency.

Can ordinal data ever be averaged?

Strictly speaking, no. Ordinal data preserves order but not equal intervals. The difference between “Agree” and “Strongly Agree” is not necessarily the same as between “Neutral” and “Agree.” Computing a mean assumes equal spacing, which is rarely justified. The median is acceptable because it only uses order. For most purposes, reporting the full distribution of responses is more informative than any single summary statistic.

What is Simpson’s paradox and how does it relate to averaging?

Simpson’s paradox occurs when a trend that appears in several groups reverses when the groups are combined. This happens because the groups have different sizes and different baseline rates. The aggregate average weights the groups unequally and can produce a result that contradicts each group’s individual trend. It demonstrates that averaging across heterogeneous categories without accounting for composition can invert the truth.

How can I tell if an average in a news report is trustworthy?

Look for three things: whether the type of average (mean or median) is specified, whether the distribution is described or shown, and whether subgroups are discussed. A trustworthy report will mention if the data is skewed, will provide the median alongside the mean, and will break down results by relevant categories. If only a single average is given with no context, the number should be viewed with skepticism.

The Arithmetic of Misunderstanding: Why Averaging the Wrong Data Leads Us Astray

Numbers feel clean. A single figure—an average, a mean, a typical value—promises to cut through the noise and give us something solid to hold onto. But that promise is often a trap. When we average data that was never meant to be averaged, we don’t simplify reality; we replace it with a tidy fiction. Jerome Leland has spent years watching how information gets packaged for public consumption, and he’s noticed a pattern: the most dangerous mistakes aren’t the obvious falsehoods. They’re the ones that look perfectly reasonable on a spreadsheet.

To see why averaging can go so wrong, you have to understand what an average actually assumes. The arithmetic mean—add everything up, divide by the count—only makes sense if the items you’re adding are fundamentally alike, members of a single coherent category. When they’re not, the average becomes a number that describes nothing real. It’s a statistical ghost, and we keep mistaking it for a living thing.

The Flaw of the “Average Person”

Back in 1945, the Cleveland Health Museum held a contest to find a woman whose body measurements matched the “average American woman” as defined by a recent anthropometric study. They found a theater cashier named Martha Skidmore and crowned her the statistical ideal. The only problem: nobody had checked whether a real person could actually match all those average dimensions at once. When the U.S. Air Force later tried designing cockpits around the average pilot’s measurements, they discovered that out of thousands of pilots, not a single one fit the average across even a handful of key dimensions. The average was a composite that described no actual human being. This is the first way averaging fails: the center point of a multivariate distribution may not correspond to any real instance in the population.

We still make this mistake today, especially in media narratives. The “average household” or the “average voter” is invoked as if such a creature exists. But households vary wildly in size, composition, and income. The average voter is a statistical ghost, smoothing over deep fractures in priorities and values. When a news story says “the average American thinks X,” it’s often papering over the very divisions that are the real story.

A diverse group of people standing together, illustrating the danger of reducing varied individuals to a single average
Reducing a diverse population to a single average figure often obscures more than it reveals.

Mixing Categories That Should Stay Separate

A second, sneakier problem shows up when we average across categories that are fundamentally different. Picture a county reporting its “average emergency response time.” The county has a dense urban core and a sprawling rural hinterland. In the city, response times clock around 6 minutes. Out in the countryside, it’s more like 25. The combined average might land at 12 minutes. That 12-minute figure is arithmetically correct, but it’s a lie in practice—it describes neither the urban nor the rural experience. It pretends two different systems are one.

You see this same error all over education reporting. A school district announces an average test score that looks respectable, but that single number buries enormous variation between schools, between demographic groups, between programs. The average becomes a tool for hiding inequality rather than exposing it. When journalists repeat these averages without unpacking them, they’re not reporting—they’re helping with the cover-up.

The Ratio Trap

Another classic case of illegitimate averaging involves ratios and rates. Imagine two hospitals. The small one performs 10 procedures and loses 1 patient—a 10% mortality rate. The large one performs 100 procedures and loses 8—an 8% rate. If you naively average the two rates, you get 9%. But the combined mortality across both hospitals is 9 deaths out of 110 procedures, or about 8.2%. The simple average of rates gives equal weight to each hospital, ignoring how many patients they actually treated. This is Simpson’s paradox in action, and it can flip the apparent truth on its head.

In public health reporting, this mistake has real consequences. Averaging infection rates across regions with vastly different population sizes can make an outbreak look milder than it is—or more severe. The right approach weights the rates by the underlying populations, but even then, the aggregate number can hide dangerous local hotspots that need attention now.

The Time-Series Deception

Averaging over time introduces its own brand of distortion. Stock market commentators love to cite average annual returns: “The market has returned an average of 7% per year over the last century.” That smooth 7% line erases crashes, booms, and long grinding periods of stagnation. An investor who entered the market at the wrong moment lived through something nothing like that smooth average. The average hides the sequence of returns, and the sequence is the actual experience of any real participant.

This temporal averaging is especially nasty in climate and weather reporting. A global average temperature rise of 1.5°C sounds almost gentle—until you realize it masks regional extremes where the rise is 3°C or 4°C, with devastating local effects. The average numbs us to the variance, and the variance is where the damage lives.

A thermometer showing a moderate temperature, while the background hints at extreme weather conditions
A single average temperature can conceal dangerous extremes that are the real story.

The Ecological Fallacy in News Narratives

When journalists report on demographic trends, they often stumble into the ecological fallacy: inferring individual behavior from group averages. A classic example is stating that counties with higher average incomes voted a certain way, and then implying that wealthier individuals voted that way. The average of a group doesn’t describe any specific member of that group, and the correlation at the aggregate level can reverse at the individual level. This error is so baked into election coverage that it’s become a structural flaw in how we understand political behavior.

Consider a report that “neighborhoods with higher average education levels have lower crime rates.” The implication is that more educated people commit fewer crimes. But the data is at the neighborhood level, not the individual level. It could be that neighborhoods with higher education levels also have more policing, better lighting, or younger populations. The average is a property of the group, not a property of the people in it. Drawing conclusions about individuals from group averages is a leap that often lands in falsehood.

When the Distribution Matters More Than the Center

In many cases, the average isn’t just misleading—it’s beside the point. Income data is the textbook example. The mean household income in the United States gets yanked upward by a small number of extremely high earners. The median, the 50th percentile, sits substantially lower. But even the median fails to capture the shape of the distribution: the clustering at the bottom, the long tail at the top, the hollowing out of the middle. Reporting only the mean or median income gives no sense of inequality, precarity, or the lived experience of most households.

Journalists who want to convey economic reality should reach for percentile ranges, Gini coefficients, or simply tell stories of specific households at different points on the curve. The average is a shortcut that short-circuits understanding.

Ordinal Data and the Meaningless Mean

Some data cannot be averaged because the numbers themselves are labels, not quantities. Survey questions that ask respondents to rate something on a scale of 1 to 5 produce ordinal data. The difference between “strongly disagree” (1) and “disagree” (2) is not the same as the difference between “agree” (4) and “strongly agree” (5). Yet it’s standard practice to compute an average rating, say 3.7, and treat it as a meaningful measure of sentiment. A 3.7 average could mean that most people were lukewarm, or that the population was polarized between 1s and 5s. The average erases that distinction.

In product reviews, movie ratings, and opinion polls, the average star rating is everywhere. It’s also frequently useless. A film with an average rating of 3 stars might be a bland, universally mediocre experience, or it might be a divisive masterpiece that half the audience adored and half despised. The latter is far more interesting, but the average hides it.

A row of five stars with the middle three highlighted, representing the ambiguity of average ratings
An average rating of 3 stars can mean consensus mediocrity or passionate division—two very different realities.

Practical Guidelines for Critical Consumption

Given how pervasive these errors are, a careful reader or reporter needs a mental checklist when confronted with an average. First, ask: What is the underlying distribution? If the data is bimodal, heavily skewed, or has extreme outliers, the average is likely deceptive. Second, ask: Are the categories being mixed? If the average combines apples and oranges—urban and rural, small and large, young and old—it probably should not exist. Third, ask: Is the average being used to make claims about individuals? If so, the ecological fallacy may be in play.

For reporters, the obligation is greater. Presenting an average without context is a form of misinformation. Whenever possible, show the distribution. Use a histogram, a box plot, or even a simple range. If you must report an average, pair it with a measure of spread—standard deviation, interquartile range, or at minimum the minimum and maximum. Let the audience see the shape of the data, not just a single point that may be a statistical mirage.

When Averages Work

This is not an argument against averages altogether. When data is normally distributed and the categories are homogeneous, the mean is a powerful and legitimate summary. If you measure the heights of a thousand randomly selected adult men from the same population, the average height is meaningful and stable. The problem arises when we lazily apply the same tool to data that violates its assumptions. The average is a precision instrument for a specific type of data, not a universal solvent.

In newsrooms, the pressure to simplify is immense. A single number fits in a headline, a chyron, a tweet. A distribution requires explanation, visualization, and nuance. But the cost of that simplification is often accuracy. The public is left with a number that feels true but is actually a distortion. Over time, these distortions accumulate into a fog of misunderstanding that shapes policy, investment, and personal decisions.

FAQ

Why is the average sometimes called a “summary statistic” if it can be so misleading?

The average is a summary statistic in the sense that it reduces many data points to a single number. This is useful when the data is symmetric and unimodal, because the average then represents a typical value. But when the data is skewed, multimodal, or contains outliers, the average summarizes poorly—it gives a number that is not typical of anything in the dataset. The term “summary” does not guarantee accuracy; it only describes the operation of condensing information.

How can I tell if an average in a news article is trustworthy?

Look for accompanying information about the distribution. A trustworthy report will mention the range, the median, or the presence of outliers. If the article only gives an average without any sense of variation, be skeptical. Also consider the source: is the average being used by an advocate to make a point, or by a disinterested analyst to describe a phenomenon? Advocacy often selects the average that supports its case while ignoring the distribution that would undermine it.

What is the difference between mean and median, and when should I prefer one over the other?

The mean is the arithmetic average: sum divided by count. The median is the middle value when data is sorted. The median is resistant to outliers and skewed distributions; it always represents a value that actually occurs in the dataset (or is halfway between two actual values). For income, home prices, and any data with a long tail, the median is usually more informative than the mean. The mean is appropriate for symmetric distributions without extreme values, such as measurement errors or natural physical traits in a homogeneous group.

Can averaging ever be completely avoided in journalism?

Probably not, and it shouldn’t be. There are legitimate uses. The key is to avoid averaging across inappropriate categories and to always provide context about the distribution. A responsible journalist treats the average as a starting point for exploration, not as the final word. When the average is the only number reported, it often becomes a stopping point for thought rather than a starting point.

The Hidden Danger of Averaging Data That Shouldn’t Be Averaged

We’ve all done it. You’ve got a column of numbers, someone asks for a quick summary, and you reach for the arithmetic mean. It’s the default move—add everything up, divide by the count, and call it a day. But here’s the uncomfortable truth: that simple average can lie to you. Not just fudge the details a little, but produce a number that points you in exactly the wrong direction. The problem isn’t the math; it’s that we keep applying it to data that was never meant to be averaged in the first place.

This isn’t a new discovery. Statisticians have been waving red flags about it for decades. Yet the habit persists, partly because the mean feels so intuitive. It’s clean, it’s quick, and it gives you a single story to tell. But when you force that story onto the wrong kind of data, the result isn’t just a harmless simplification—it’s a quiet disaster waiting to unfold in a budget decision, a medical trial, or a policy that affects millions.

The Nature of the Problem: When Numbers Aren’t Just Numbers

To see why averaging can go so wrong, you have to think about what an average actually does. It assumes each value is a standalone, additive piece of the puzzle. If you’re measuring the weight of ten watermelons, the average weight makes perfect sense. Each watermelon contributes its own mass, and the sum divided by ten gives you a solid idea of what a typical melon weighs. No drama there.

But a lot of the numbers we deal with aren’t like watermelon weights. They’re ratios, percentages, or rates that have their own internal arithmetic—often with hidden denominators that vary wildly. When you average them directly, you’re pretending those denominators don’t exist. You’re treating a percentage based on 10 observations as if it carries the same weight as one based on 10,000. That’s not just sloppy; it’s a recipe for conclusions that are the opposite of reality.

Person analyzing data on a laptop with charts and graphs on screen

The Classic Trap: Average of Averages

Let’s make this concrete. Say you’re looking at customer satisfaction for three products. Product A has 100 reviews and a 4.5 average. Product B has 10 reviews and a 4.0 average. Product C has 5 reviews and a 2.0 average. If you just average those three numbers—(4.5 + 4.0 + 2.0) / 3—you get 3.5. Sounds reasonable, right? But it’s almost meaningless. That 3.5 treats the 5 reviews for Product C as equal to the 100 reviews for Product A. The real overall satisfaction across all 115 reviews is (100×4.5 + 10×4.0 + 5×2.0) / 115 = 4.35. The simple average of averages understates the truth by nearly a full point.

This isn’t a classroom hypothetical. I’ve seen it in business reports, school district summaries, and public health dashboards. A district might report the average pass rate across its schools as a simple mean, quietly burying the fact that a large school with a high pass rate is being drowned out by a tiny school with a low one. The number they publish doesn’t reflect what most students actually experience. It’s a statistical phantom.

When Percentages Go Rogue

Percentages pull the same trick. Picture a company tracking conversion rates for two campaigns. Campaign A draws 1,000 visitors and converts at 5%. Campaign B draws 10 visitors and converts at 50%. Average those two percentages and you get 27.5%. That’s a number that would make any marketing team pop champagne. But the true overall conversion rate is (50 + 5) / (1000 + 10) = 55 / 1010 ≈ 5.45%. The 27.5% figure is a fiction, inflated by giving equal weight to a tiny, fluky campaign. Reporting that “average” would set expectations the business could never meet.

Medicine has its own version of this headache. When researchers compare success rates of treatments across hospitals of different sizes, averaging the percentages without weighting by patient count can flip the conclusion entirely. This is Simpson’s paradox in action—a trend that appears in the aggregated data reverses when you properly disaggregate or weight the groups. It’s not just a statistical curiosity; it has steered clinical decisions and health policies in the wrong direction.

Close-up of a calculator and financial reports on a desk

When the Process Is Multiplicative, Not Additive

Another common stumble is using the arithmetic mean for data that grows or shrinks through compounding—investment returns, population growth, bacterial expansion. Here, the arithmetic mean overstates the typical outcome because it ignores how gains and losses build on each other.

Take an investment that jumps 50% in year one and then drops 50% in year two. The arithmetic mean of those returns is a calm 0%. But start with $100. After year one, you’ve got $150. After year two, you’re down to $75. That’s a 25% loss, not a break-even. The right tool here is the geometric mean, which respects compounding. The geometric mean of 1.5 and 0.5 is √(1.5 × 0.5) = √0.75 ≈ 0.866, an average loss of about 13.4% per year. That’s what actually happened to the money.

In finance, this mistake feeds overly rosy projections. In biology, it warps growth rate estimates. Anywhere values multiply rather than add, the arithmetic mean is simply the wrong wrench for the job.

When the Data Lopsides

The arithmetic mean also falls apart with skewed distributions. Income data is the poster child. A handful of ultra-high earners yank the mean income far above what a typical person sees. The median—the point where half earn more and half earn less—gives a much straighter picture. Yet news reports and political speeches keep quoting mean income, painting a distorted portrait of economic health.

This skew problem shows up anywhere outliers can hijack the average. In website performance monitoring, a few agonizingly slow page loads can inflate the mean load time, hiding the fact that most users have a snappy experience. That’s why engineers lean on percentiles—the median or the 95th percentile—to understand both typical and worst-case behavior. Averaging skewed data without acknowledging the skew is a fast track to bad decisions.

The Ecological Fallacy: Averages Across Groups

Then there’s the ecological fallacy—drawing conclusions about individuals from group-level averages. If you hear that a neighborhood’s average income is $80,000, you can’t assume a random resident earns about that much. The neighborhood might be a patchwork of very wealthy and very poor households, with few in the middle. The average smooths over that reality.

This fallacy creeps into public policy constantly. Crime rates averaged across large geographic areas can mask intense hotspots. Average test scores for a school can hide wide gaps between student subgroups. Whenever an average is used to make claims about individuals or smaller units, you have to dig into the variability underneath. Otherwise, the “average” becomes a convenient fiction that misdirects resources and interventions.

Person pointing at data on a whiteboard with graphs and charts

When the Data Points Aren’t Independent

Averaging also assumes each data point is an independent, equal voice. But in many real-world datasets, observations are clustered or correlated. In education research, students are nested in classrooms, which are nested in schools. Averaging test scores across all students without accounting for that structure can underestimate variability and lead to wrong conclusions about what works. Hierarchical models are built for this, but the simple average remains a common—and flawed—shortcut.

Time-series data has a similar issue. Consecutive observations are often correlated. Averaging daily stock prices over a month ignores the volatility and trends within the month. The average temperature for a year says nothing about extreme heat waves or cold snaps. In these cases, the average steamrolls over the very patterns that matter most.

When the Average Isn’t Even Possible

Sometimes the arithmetic mean spits out a number that can’t exist in the real world. The old joke about the average family having 2.4 children is harmless at a barbecue, but it’s less funny in planning. If a school district uses an average class size of 22.7 to allocate teachers, some classes will end up far above or below that number, creating real inequities. The average hides the distribution, and the distribution is what affects actual people.

In manufacturing, an average defect rate of 0.2% might sound acceptable. But if defects cluster in specific batches, that average conceals a serious quality control problem. The mean is not a stand-in for understanding variation.

Practical Questions to Ask Before You Average

So how do you know when averaging is safe and when it’s a trap? Here are the questions I run through before computing a mean:

  • Are the values additive? If adding them up doesn’t make physical or logical sense, the arithmetic mean is suspect. For ratios and rates, think weighted means or harmonic means.
  • Is the underlying process multiplicative? For growth rates, returns, or anything that compounds, use the geometric mean.
  • Is the distribution highly skewed? If so, the median or a trimmed mean may be more representative.
  • Are the denominators equal? When averaging percentages or rates, check that the base counts are similar, or weight appropriately.
  • Are the data points independent? If there’s clustering or correlation, simple averaging can mislead.
  • Does the average correspond to a real, possible value? If not, consider whether the mean is the right summary at all.

These aren’t academic nitpicks. They have real-world consequences. In 2013, an analysis of clinical trial data for a blood thinner showed that averaging results across different patient subgroups had masked a dangerous increase in bleeding risk for certain patients. The drug had been approved based on the overall average, which looked safe. Proper subgroup analysis told a different, more alarming story.

Why the Mistake Keeps Happening

Despite decades of warnings, the misuse of averages is everywhere. Part of it is that averages are easy to compute and easy to explain. They boil complexity down to a single number, which is seductive in a world that craves simplicity. But another part is that many people who work with data—managers, journalists, policy analysts—never got formal training in statistical reasoning. They reach for the mean because it’s the only tool on their belt.

There’s a deeper cognitive bias at work, too. Our minds love central tendencies and get uncomfortable with dispersion. We want a single story, not a distribution of possibilities. Averages hand us that story, even when it’s false. Pushing back against that bias takes discipline and a willingness to sit with the messiness of real data.

FAQ

Why is averaging percentages often misleading?

Averaging percentages directly ignores differences in the base sizes (denominators) from which those percentages were calculated. A percentage from a small sample can have an outsized influence on the simple average, distorting the true overall rate. The correct approach is to weight each percentage by its denominator or to compute the percentage from the combined totals.

When should I use the median instead of the mean?

The median is preferable when the data distribution is skewed or contains extreme outliers. Unlike the mean, the median is not pulled toward the tail of the distribution, so it better represents a “typical” value. Common examples include income, house prices, and reaction times.

What is Simpson’s paradox, and how does it relate to averaging?

Simpson’s paradox occurs when a trend that appears in several different groups of data disappears or reverses when the groups are combined. It often arises from averaging rates or percentages across groups with very different sizes, leading to a summary statistic that contradicts the patterns within each group. Proper weighting or stratified analysis is required to avoid the paradox.

Why can’t I just average growth rates over time?

Growth rates compound multiplicatively, not additively. The arithmetic mean of a series of growth rates ignores the compounding effect and will generally overstate the true average growth. The geometric mean correctly accounts for compounding and provides the constant rate that would yield the same final value over the period.

In the end, the problem with averaging data that should not be averaged isn’t a technical quirk—it’s a fundamental failure to respect the structure of the information. The arithmetic mean is a model, and like any model, it has assumptions. When those assumptions are violated, the model breaks. The responsible analyst knows when to use the mean, when to reach for an alternative, and when to simply present the data in its full, unvarnished complexity. Because sometimes, the most honest summary is not a single number at all.

The Hidden Danger of Averaging Data That Should Never Be Averaged

Numbers have a way of seducing us. They promise clarity, precision, a neat path through the mess. But there’s a quiet, stubborn error that slips into reports, dashboards, and news headlines more often than anyone wants to admit: averaging data that, by its very nature, should never be averaged. Jerome Leland has watched this mistake play out across finance, healthcare, public policy—and the fallout is rarely small.

The urge to average makes sense on the surface. A single figure feels clean. It slides into a headline, a bullet point, a ten-second sound bite without friction. But when the underlying data is built from ratios, percentages of mismatched wholes, or distributions that lean hard in one direction, the resulting “average” isn’t just misleading. It’s mathematically incoherent. It conjures a number that never actually existed in any real-world observation.

The Arithmetic Trap of Averages

At its heart, an average is a measure of central tendency. The arithmetic mean—sum everything up, divide by the count—works beautifully for additive quantities. Heights of ten people? Add them, divide by ten, and you get a meaningful central height. Same goes for weights, temperatures, or test scores where each observation stands on its own, measured on the same scale.

Trouble starts when the data points are themselves ratios, rates, or percentages drawn from different bases. Picture a small hospital that handles 10 surgeries and sees 1 complication—a 10% complication rate. A large hospital handles 1,000 surgeries and logs 50 complications, a 5% rate. Average 10% and 5% naively, and you get 7.5%. But the combined reality across both hospitals is 51 complications out of 1,010 surgeries—a rate of 5.05%. That 7.5% figure is a phantom. It describes no actual population of patients.

This error has a dramatic cousin called Simpson’s Paradox, where trends in subgroups flip when data gets aggregated improperly. But the underlying problem is simpler: averaging ratios without weighting them by their denominators produces a number detached from reality. Yet the practice is everywhere—in business metrics, educational statistics, even medical research summaries.

When Percentages Collide

Think about a marketing analyst comparing email open rates across five campaigns. Each campaign hits a different audience size. Campaign A reaches 100,000 recipients with a 20% open rate. Campaign B reaches 1,000 recipients with an 80% open rate. The simple average of 20% and 80% is 50%. But the actual combined open rate across both campaigns isn’t 50%—it’s yanked heavily toward the larger campaign. The true overall rate is (20,000 + 800) / (100,000 + 1,000) = 20.6%. Reporting 50% would be a catastrophic overstatement.

This isn’t academic. Jerome Leland has sat through investor pitch decks where customer retention rates were averaged across cohorts of wildly different sizes, producing a figure that suggested far healthier retention than actually existed. The startup’s largest cohort had a 40% retention rate, while a tiny pilot cohort showed 90%. The simple average of 65% masked the fact that the vast majority of customers were churning at a much higher rate.

Business professionals reviewing charts and data on a whiteboard
Data presentations often hide the denominator disparities that make simple averages misleading.

The Denominator Problem

At the bottom of this statistical misstep is a failure to respect the denominator. Every percentage, rate, or ratio is a fraction: a numerator divided by a denominator. The denominator represents the scale, the base, the population from which the numerator is drawn. When you average percentages without accounting for their denominators, you implicitly assume all denominators are equal. That assumption is rarely true and often wildly false.

Take crime rates across cities. City A has a population of 100,000 and reports 500 violent crimes, a rate of 5 per 1,000 residents. City B has a population of 10,000 and reports 100 violent crimes, a rate of 10 per 1,000. The simple average of the two rates is 7.5 per 1,000. But if you combine the cities, the actual rate is 600 crimes per 110,000 residents, or about 5.45 per 1,000. The unweighted average overstates the true rate by nearly 40%.

This distortion gets worse as the gap between denominators widens. When a small group with an extreme rate is averaged equally with a large group with a moderate rate, the small group punches above its weight. The result is a number that’s mathematically possible but empirically false—a synthetic statistic that represents no actual population.

Skewed Distributions and the Mean’s Deception

Another domain where averaging falls apart is highly skewed data. Income distributions, for example, are famously right-skewed: a small number of extremely high earners pull the arithmetic mean far above the median. Reporting the “average income” in such cases gives a distorted picture of what a typical person actually earns. The mean becomes a measure not of centrality but of total wealth divided by population, which is a different concept entirely.

In 2019, the U.S. Census Bureau reported a median household income of $68,703, while the mean household income was $98,088. That $30,000 gap isn’t a rounding error; it’s the gravitational pull of the top percentiles. When journalists or analysts casually refer to the “average American household income” without specifying which average, they often inadvertently misrepresent the economic reality for most households.

This problem extends to any metric with a long-tailed distribution: website page load times, hospital wait times, customer lifetime value, or the size of insurance claims. In each case, a few extreme values can inflate the arithmetic mean to a point where it no longer describes a typical experience. The median or geometric mean often provides a more honest central value, yet the arithmetic mean persists because it’s computationally simple and familiar.

When Averaging Destroys Information

Beyond the mathematical errors, there’s a deeper conceptual problem: averaging can destroy the very information that matters most. In public health, averaging infection rates across regions with vastly different population densities and healthcare access can hide dangerous hotspots. A national average might look reassuring while a specific city is experiencing an outbreak that requires immediate intervention.

In education, averaging test scores across schools with different demographics, funding levels, and class sizes can mask severe inequities. A district-wide average might suggest acceptable performance while individual schools are failing. The average becomes a tool of erasure, smoothing over the variation that policymakers most need to see.

In finance, averaging portfolio returns across time periods without considering volatility or sequence of returns can lead to disastrous planning. A retiree withdrawing funds during a market downturn faces a fundamentally different reality than the “average annual return” would suggest. The average hides the sequence-of-returns risk that can destroy a retirement plan.

Person analyzing financial charts and graphs on a laptop screen
Financial averages often conceal the volatility and sequence risk that determine real-world outcomes.

The Simpson’s Paradox Trap

Simpson’s Paradox is the most famous manifestation of the averaging error, and it deserves a closer look. The paradox occurs when a trend appears in several different groups of data but disappears or reverses when the groups are combined. The classic example comes from a 1973 study of gender bias in graduate admissions at the University of California, Berkeley. When looking at aggregate data, men appeared to be admitted at a higher rate than women. But when the data was broken down by department, women had equal or higher admission rates in most departments. The paradox arose because women tended to apply to more competitive departments with lower overall admission rates.

This is not a statistical curiosity; it’s a recurring trap in business analytics. A company might see that its overall customer satisfaction score is declining, yet every individual product line shows improving satisfaction. The paradox occurs because customers are migrating toward a product line with historically lower satisfaction scores. The aggregate trend points in the opposite direction of every subgroup trend. Acting on the aggregate number alone would lead to completely wrong conclusions.

The lesson isn’t that averages are useless, but that they must be applied with an understanding of the underlying data structure. When subgroups have different sizes and different base rates, the simple average is not just imprecise—it can be actively deceptive.

When Averaging Time-Series Data Misleads

Another common error is averaging data across time periods that have fundamentally different characteristics. Consider a retailer reporting “average monthly sales” over a year that includes both normal months and the holiday season. The average will be pulled upward by November and December, creating a figure that represents no actual month. A manager using this average to set monthly targets will find it too high for most of the year and too low for the peak season.

Similarly, averaging economic indicators across business cycle phases—recessions and expansions—produces a number that describes neither. The average GDP growth rate over a period that includes both a boom and a bust is a statistical artifact, not a meaningful benchmark. Yet such averages routinely appear in economic commentary as if they represented a stable underlying trend.

In medicine, averaging patient outcomes across different risk groups can obscure critical safety signals. A drug might show acceptable average efficacy, but when the data is stratified by age or comorbidity, it becomes clear that the drug is highly effective for some patients and dangerous for others. The average hides this heterogeneity, potentially leading to harmful prescribing practices.

Alternatives to the Naive Average

When faced with data that should not be simply averaged, several alternatives exist. The choice depends on what question is actually being asked.

Weighted averages correct the denominator problem by giving each data point influence proportional to its size. For the hospital complication rates, weighting by the number of surgeries at each hospital produces the correct aggregate rate. This is the most direct fix for the ratio-averaging error and should be the default approach whenever combining rates or percentages from groups of different sizes.

Medians resist the pull of extreme values and better represent the “typical” observation in skewed distributions. For income data, home prices, or any metric with a long tail, the median is almost always more informative than the mean for describing what a typical person or case experiences.

Geometric means are appropriate for averaging growth rates, returns, or other multiplicative processes. If an investment grows 50% one year and falls 50% the next, the arithmetic mean return is 0%, but the geometric mean correctly shows a net loss. The arithmetic mean of percentage changes over time is almost always misleading.

Stratified reporting preserves the variation that simple averaging destroys. Instead of reporting a single average, report the averages for each meaningful subgroup alongside the overall weighted average. This approach maintains transparency and allows readers to see both the forest and the trees.

Multiple charts and graphs displayed on a wall showing different data segments
Stratified reporting preserves the variation that simple averages erase.

Real-World Consequences

The consequences of averaging data that should not be averaged extend far beyond academic debates. In 2016, a major pharmaceutical company presented aggregated clinical trial results that showed a modest benefit for a new drug. Only when independent researchers reanalyzed the data by patient subgroups did it become clear that the drug was actively harmful to patients over 65 while providing significant benefit to younger patients. The simple average had masked a critical safety signal.

In public policy, averaging unemployment rates across regions with vastly different economic structures can lead to misallocation of resources. A national unemployment rate of 5% might sound healthy, but if it consists of 2% unemployment in urban areas and 15% in rural communities, the average hides a crisis. Programs designed based on the national average will fail to address the real problem.

In education, the No Child Left Behind Act required schools to report average test scores, which led to well-documented instances of schools focusing resources on students near the proficiency threshold while neglecting both high-achieving and severely struggling students. The average became the target, and the target distorted the system.

How to Spot the Error

Readers and analysts can protect themselves by asking a few diagnostic questions whenever they encounter an average:

First, what is being averaged? If the answer is a percentage, rate, or ratio, immediately check whether the denominators are equal. If they are not, the average is almost certainly wrong unless it has been explicitly weighted.

Second, what is the distribution? If the data is skewed, the mean will be pulled away from the typical value. Ask for the median or a full distribution plot. A single number cannot capture a skewed reality.

Third, what variation is being hidden? Even a correctly calculated weighted average can conceal important differences between subgroups. Always ask for the disaggregated data. If the source cannot or will not provide it, treat the average with skepticism.

Fourth, is the average being used to make a decision that depends on the full distribution? In risk management, inventory planning, or capacity sizing, the average is often insufficient. The extremes—the tails of the distribution—frequently determine success or failure. A bridge designed to withstand the average flood will collapse regularly.

FAQ

Why can’t I just average percentages from different groups?

Percentages are ratios with denominators that often differ in size. Averaging them without weighting by those denominators gives equal influence to groups of unequal size, producing a number that does not correspond to the actual combined rate. For example, averaging a 90% rate from a group of 10 and a 40% rate from a group of 1,000 yields 65%, but the true combined rate is about 40.5%. The simple average is mathematically incoherent because it ignores the scale of each group.

When is it acceptable to use a simple average?

A simple arithmetic mean is appropriate when the data points are additive quantities measured on the same scale and each observation carries equal weight. Examples include averaging the heights of individuals in a room, the temperatures recorded at different weather stations on the same day, or the scores of students who all took the same test. The key condition is that each data point represents a measurement of the same kind with a consistent underlying unit.

What should I use instead of a simple average for skewed data?

For skewed distributions, the median is usually a better measure of central tendency because it is not pulled by extreme values. For growth rates or returns over time, the geometric mean is more appropriate because it accounts for compounding effects. For data with unequal group sizes, a weighted average that uses the group sizes as weights will give the correct aggregate figure. In many cases, reporting the full distribution or providing stratified averages alongside the overall figure is the most honest approach.

How does averaging cause bad business decisions?

Averaging can lead to bad decisions when it hides variation that matters for planning. If a retailer uses average monthly sales to set inventory levels, they will be overstocked in slow months and understocked in peak months. If an investor uses average annual returns to plan retirement withdrawals without considering the sequence of those returns, they may run out of money during an early downturn. The average provides a single number, but real-world outcomes depend on the entire pattern of variation.

Why Housing Market Data Varies Wildly by Region

National housing statistics have a way of sounding like they were written for a country that doesn’t actually exist. A headline says home prices rose 5% nationwide, and somewhere in Ohio a buyer nods along while a seller in Boise wonders what planet that number came from. In one county, prices jumped 18%. In another, they didn’t budge. This isn’t a glitch in the data. It’s what happens when local economies, land constraints, and the slow churn of demographics pull real estate in completely different directions. To make sense of it, you have to look past the tidy national averages and dig into the structural forces that turn every metro area, suburb, and rural county into its own market.

Aerial view of suburban housing development with varied home styles

The Illusion of a National Housing Market

Real estate is stubbornly local. A house in San Francisco can’t be trucked to Cleveland to meet demand there, and a construction labor shortage in Phoenix doesn’t touch the supply of homes in Atlanta. Yet the numbers that grab the most attention—the S&P CoreLogic Case-Shiller National Home Price Index, the median sales price from the National Association of Realtors—lump hundreds of distinct markets into a single figure. That’s handy for tracking broad economic currents, but it buries the wild variation underneath. In any given month, the spread of price changes across the 100 largest metro areas can run three or four times the national average change. A national index is basically a weighted average, and the weights themselves—population, transaction volume—shift over time, adding another layer of distortion for anyone trying to read a specific location.

If you want to understand why regional data pulls apart so dramatically, picture each housing market as a product of three interacting layers: the economic base that generates incomes and jobs, the physical and regulatory environment that chokes or frees supply, and the demographic profile that shapes who’s buying and what they want. When those layers stack up differently from place to place, you get a patchwork of price trajectories, inventory levels, and sales speeds that can look almost unrelated.

Layer One: The Economic Engine

Job growth is the bluntest instrument driving housing demand in a region. A metro area adding jobs at 3% a year while the national rate limps along at 1.5% gets an immediate surge of workers who all need somewhere to live. But the kind of jobs matters as much as the count. A market stacked with high-wage tech or finance positions reacts differently than one adding mostly retail or hospitality jobs. In the Bay Area, where the median tech salary tops $150,000, even a modest hiring bump at a few big firms can set off bidding wars that push prices in certain ZIP codes up by double digits in a single quarter. Meanwhile, a manufacturing hub in the Midwest might add thousands of jobs at $55,000 a year and see only a gentle price uptick because the income-to-price ratio stays more grounded.

Economic diversification—or the lack of it—also leaves a mark. Regions that lean heavily on a single industry—oil extraction in West Texas, tourism in Orlando, government in Washington, D.C.—ride housing cycles that are lashed to the fortunes of that sector. When oil prices cratered in 2014–2015, housing markets in Midland and Odessa dropped 10–15% while the rest of the country was still climbing out of the Great Recession. On the flip side, the remote-work shift that took off in 2020 dropped high-earning professionals into markets like Boise, Idaho, and Bend, Oregon, where local wages had never propped up such price levels. The result was a sudden, sharp re-pricing that national indices barely noticed until the trend had been running for months.

Modern downtown office buildings reflecting economic activity

Layer Two: The Supply Straitjacket

If demand is the accelerator, supply is the brake—and in a lot of regions, the brake is rusted solid. The ability to add housing when demand picks up varies wildly by geography and regulation. Broadly, markets fall into three supply buckets: elastic, constrained, and severely restricted.

Elastic Markets: The Plains and the Sunbelt Periphery

Across much of Texas, the Great Plains, and the outer edges of Sunbelt metros, land is plentiful and zoning is loose. When demand rises, developers can grab large tracts, pull permits, and deliver new subdivisions or apartment complexes in 18 to 24 months. That elasticity keeps price growth on a leash. Houston, for example, has absorbed more than a million new residents per decade while keeping median home prices comfortably below the national average. The data from these places often shows climbing sales volume and prices that are stable or rising slowly—a pattern that looks almost sleepy next to the coastal spikes.

Constrained Markets: Established Suburbs and Mid-Size Cities

Many older suburbs and mid-size cities sit in a middle ground. Land is still out there, but it needs pricier infrastructure extensions, and local approval processes add time and cost. These markets can respond to demand, but with a lag and at a higher price floor. The data here tends to follow a cyclical rhythm: prices climb during demand surges, then flatten or dip a little as new supply catches up. Markets like Nashville, Charlotte, and Raleigh-Durham fit this pattern—price growth above the national average but not explosive, inventory levels that oscillate inside a predictable band.

Severely Restricted Markets: Coastal Gateways and Land-Locked Cities

At the far end are markets where physical geography, tight zoning, and political resistance combine to make new construction a slow-motion ordeal. San Francisco, Los Angeles, Boston, New York, and Seattle all share this profile. In these places, the supply curve is nearly vertical: even a massive demand surge produces only a trickle of new units. Almost all the demand pressure converts straight into higher prices. When mortgage rates drop, these markets don’t see a jump in sales volume; they see a jump in bidding wars and price per square foot. The data is marked by low inventory, high price volatility, and a strong inverse relationship between rates and prices—a relationship that’s much weaker in elastic markets.

Regulatory differences deepen the physical constraints. California’s Environmental Quality Act (CEQA) and local discretionary review processes can stretch project timelines by years. In contrast, Texas’s minimal zoning and by-right development rules let projects move on predictable schedules. Those policy choices are baked into the housing data: markets with more regulatory friction show higher price sensitivity to demand shocks and a weaker response to supply-side fixes.

Layer Three: Demographic Currents

Population growth alone doesn’t set housing demand; the makeup of that growth matters just as much. Two regions adding 50,000 households a year can end up with completely different housing outcomes depending on whether those households are young renters, families hunting single-family homes, or retirees downsizing.

Millennials, now the largest homebuying group, are forming households at a fast clip, but their geographic preferences are uneven. They’re clustering in Sunbelt metros with strong job markets and relative affordability—Austin, Denver, Tampa. That concentration supercharges demand in those specific markets while leaving others with weaker demographic tailwinds. Meanwhile, baby boomers are aging in place at high rates in the Northeast and Midwest, choking turnover and keeping inventory tight even where population growth is flat. The data reflects this: markets with high boomer homeownership rates often show low sales volume but steady prices, because the few listings that appear get snapped up by the buyers who can still afford them.

International migration adds another layer of regional texture. Gateway cities like Miami, Los Angeles, and New York pull in disproportionate shares of foreign-born residents, many of whom bring capital for home purchases. That inflow can prop up price levels even when domestic demand softens. During stretches of global economic uncertainty, these markets can decouple from national trends entirely, driven by currency swings and geopolitical events that have no bearing on, say, the housing market in Kansas City.

Diverse group of people walking in a residential neighborhood

How the Data Itself Creates Divergence

Even when underlying market conditions are similar, the measurement of housing data can spit out readings that don’t match. Different data sources track different slices of the market, use different methods, and operate on different clocks. The median sales price reported by a local Realtor association reflects only homes sold through the Multiple Listing Service (MLS) during a specific window. It can get yanked around by a shift in the mix of homes sold: if more luxury homes close in a given month, the median rises even if no individual home appreciated. This compositional effect hits small markets especially hard, where a handful of high-end deals can swing the median by several percentage points.

Repeat-sales indices, like the Case-Shiller tiered indices, control for compositional bias by tracking price changes on the same properties over time. But these indices cover only a subset of markets—usually the largest metro areas—and exclude new construction, which can understate price trends in fast-growing regions where new homes make up a big share of transactions. Appraisal-based measures, used by the Federal Housing Finance Agency (FHFA), pull from conforming loan data and miss jumbo loans and cash deals that dominate high-cost markets. Every methodology has blind spots, and those blind spots line up differently with regional market structures.

Timeliness also varies. MLS data is available within days of a sale closing. Case-Shiller indices publish with a two-month lag. In a fast-moving market, that lag can make the index look disconnected from what’s happening on the ground. During the early months of the COVID-19 pandemic, real-time listing data showed sharp price increases in suburban and rural markets, but the official indices didn’t confirm the trend until late summer 2020. By then, the narrative had already shifted, and the data seemed to be catching up rather than leading.

Interest Rates: A National Lever with Local Consequences

Mortgage rates are set in national and global financial markets, so they apply evenly across the country. But the effect of a rate change is anything but even. In a market where the median home price is $200,000, a one-percentage-point rate hike might tack $100 onto the monthly payment—noticeable, often manageable. In a market where the median is $1.2 million, the same rate hike adds $600 or more per month, shoving a lot of buyers out of qualification entirely. This asymmetry means rate hikes clobber high-cost markets harder than low-cost ones, while rate cuts light them up more explosively.

The mix of adjustable-rate mortgages (ARMs) and cash purchases further tweaks the rate effect. In luxury markets and investor-heavy regions, cash transactions can account for 30–40% of sales, insulating those segments from rate moves. In entry-level markets dominated by FHA and VA loans, rate sensitivity runs much higher. Data from Miami or Manhattan may show price resilience during rate tightening cycles, while data from outer-ring suburbs of the same metros shows sharp slowdowns—all because the financing mix differs across price tiers and locations.

Inventory Dynamics: More Than Just Supply and Demand

Months of supply—the ratio of active listings to the monthly sales pace—is the standard gauge for market balance. A reading below four months usually signals a seller’s market; above six months points to a buyer’s market. But this metric behaves differently depending on the underlying turnover rate of the housing stock. In a high-turnover market, like a transient military town, a three-month supply might be normal and not especially tight. In a low-turnover market, like a stable suburban community where homeowners stay put for 15–20 years, a three-month supply represents genuine scarcity.

Regional differences in housing type also mess with inventory metrics. Markets with a high share of condominiums—Honolulu, Chicago’s downtown—can see inventory swings driven by investor activity and short-term rental rules rather than owner-occupant demand. A crackdown on Airbnb listings can flood the for-sale market with condos, pushing months of supply higher even as single-family home inventory stays tight. The aggregate metro figure would show a balanced market, but that number would be useless for a family searching for a detached home in a specific school district.

Policy Interventions: Local Experiments, Local Results

Housing policy in the United States is overwhelmingly local. Zoning codes, rent control ordinances, inclusionary housing requirements, and property tax structures are set by cities, counties, and states, creating a patchwork of policy experiments whose results show up in the data. Minneapolis’s 2040 plan, which eliminated single-family-only zoning citywide, has been linked to a moderation in rent growth relative to peer cities, even as home prices kept rising. Oregon’s statewide rent control law, enacted in 2019, appears to have slowed rent increases in Portland but may have also trimmed the supply of new rental units as developers shifted to for-sale construction. These effects are visible only when you compare the affected market to similar markets without the policy—a comparison that national data can’t provide.

Property tax regimes also carve out regional divergences. In high-tax states like New Jersey and Illinois, property tax bills can top $10,000 a year on a median-priced home, acting as a drag on price appreciation because buyers factor the ongoing cost into their affordability math. In low-tax states like Colorado and Alabama, the same home price carries a much lighter tax burden, letting more of the buyer’s income go toward the mortgage principal. Over time, that differential compounds, feeding faster price growth in low-tax markets even when other conditions look similar.

Climate Risk and Insurance: The Emerging Divider

A newer factor pulling regional housing data apart is the cost and availability of property insurance. In coastal Florida, parts of California, and wildfire-prone stretches of the West, insurance premiums have shot up—doubling or tripling in a few years in some spots—and some insurers have pulled out of markets entirely. This hits housing affordability directly and, by extension, demand. A home in a high-risk area may carry a lower sticker price than a comparable home in a safe area, but the total cost of ownership, insurance included, can be higher. The data doesn’t always capture that nuance: a median price decline in a fire-prone county might look like a buying opportunity, when in reality it’s the market pricing in the rising cost of risk.

Flood zone designations, wildfire risk scores, and hurricane exposure are starting to appear in some multiple listing services and automated valuation models, but the practice is spotty. Two homes with identical square footage and bedroom counts can have wildly different values based on their risk profiles, and those differences are increasingly driving price divergence within regions as well as between them.

Interpreting Regional Data: A Framework for Readers

Given all these sources of variation, how should someone reading housing market news make sense of the numbers? The first step is to identify the data source and understand its limits. A national median price from the NAR is a decent temperature check but not a diagnostic tool for any specific market. A Case-Shiller index is better for tracking price trends over time in covered metros, but it won’t capture new construction or condos. Local MLS data is the most granular and timely, but it’s subject to compositional bias and may not include every transaction.

The second step is to look at multiple indicators together. Price change alone isn’t enough; it should be viewed alongside inventory levels, days on market, the share of listings with price reductions, and building permit activity. A market with rising prices and rising inventory is probably nearing a peak. A market with rising prices and falling inventory is still undersupplied. A market with flat prices but surging permits may be about to see a correction as new supply arrives.

The third step is to break things down by price tier and property type. Many metro areas have split markets where the luxury segment is cooling while the entry-level segment is still hot, or where condos are languishing while single-family homes are appreciating. The aggregate median can hide these splits, so drilling down to the segment that matches your situation is essential.

FAQ: Understanding Regional Housing Data

Why do two neighboring counties sometimes have completely different housing market trends?

Neighboring counties can pull apart sharply because of differences in school district quality, property tax rates, zoning restrictions, and commute patterns. A county with top-rated schools and a short commute to a major employment center commands a premium that can stick even when the broader region slows. On top of that, one county may have open land for new construction while the other is built out, leading to different supply responses to the same demand shock. Local government policies—impact fees, growth boundaries, permit processing times—can widen the gap further.

How can I tell if a local market report is reliable?

Look for reports that lay out their methodology clearly. Reliable reports will name the data source (MLS, public records, or a repeat-sales index), the geographic coverage, the time period, and whether the figures are seasonally adjusted. Be wary of reports that lean only on median prices without addressing compositional shifts, or that draw broad conclusions from small sample sizes. Cross-checking with other sources—the FHFA’s all-transactions index for your metro area, local building permit data—can help confirm the trends.

Why do housing markets in some cities seem immune to interest rate increases?

Markets with a high concentration of cash buyers—luxury destinations, investor-heavy metros—are less sensitive to mortgage rate changes because a big share of transactions don’t involve financing. Also, markets with severe supply constraints may see prices hold steady even as sales volume drops, because the few buyers left are competing for a very limited pool of listings. In these markets, rate increases shrink the number of transactions more than they shrink prices, creating a misleading impression of stability.

Is it better to follow national or local housing data when making a decision?

For any individual buying or selling decision, local data matters far more than national data. National trends can offer context—whether the overall credit environment is tightening or loosening—but the specific conditions of your city, neighborhood, and price tier will shape your experience. Even within a metro area, conditions can shift by ZIP code. The most effective approach is to track local data regularly, understand its quirks, and use national data only as a backdrop for broader economic conditions.

Housing market data varies wildly by region because housing markets themselves are wildly different. The forces that shape them—jobs, land, rules, people, and risk—are spread unevenly across the country, and the data we use to measure them is imperfect and incomplete. Recognizing those sources of variation is the first step toward reading the numbers with the precision they demand.

How to Read a Federal Budget Proposal Without Getting Lost

Stacks of printed federal budget documents on a desk with a calculator and pen

Jerome Leland here. If you’ve ever opened a federal budget proposal, you know the feeling—staring at thick volumes, tables that sprawl across pages, sentences that seem built to conceal rather than explain. It’s not a single number or a tidy ledger. It’s a policy platform, a political wish list, and a fiscal blueprint all mashed together. This guide walks you through the structure, the jargon, and a few analytical habits that turn a disorienting stack of paper into something you can actually work through.

Let’s clear up one big misunderstanding right away: the President’s budget proposal isn’t law. It’s a request to Congress, usually landing in February or March. Congress treats it as a starting point—often ignoring large chunks—while drafting its own budget resolution and appropriations bills. What you’re reading is a statement of priorities, loaded with assumptions about economic growth, revenue, and spending that may or may not pan out. Grasping that distinction by itself will spare you a lot of confusion.

Where to Find the Core Documents

The President’s budget comes from the Office of Management and Budget and lives on whitehouse.gov/omb. The main book is called “Budget of the U.S. Government,” and it runs several hundred pages. Alongside it, you’ll find the “Analytical Perspectives” volume—covering crosscutting topics like infrastructure or credit programs—and the “Appendix,” with line-by-line account detail. If you want raw numbers without the policy narrative, the “Historical Tables” are your friend: they show actual outlays and receipts stretching back decades, free of projections and spin.

Start with the main budget volume. It opens with a message from the President and a summary chapter that lays out the headline numbers: total outlays, total receipts, the deficit, and the debt. This chapter is where the administration frames its story. Read it closely, but read it with skepticism. The figures you see rest on economic assumptions—GDP growth, inflation, interest rates—that can be rosy or restrained. The budget always includes a table labeled “Economic Assumptions” or something similar. Check those assumptions against outside forecasts, like the Congressional Budget Office or the Federal Reserve, to get a sense of how realistic the picture is.

Open federal budget book with charts and graphs visible

Decoding the Major Categories

Federal spending falls into three broad buckets: mandatory spending, discretionary spending, and net interest on the debt. If these categories aren’t clear to you, the rest of the document will stay foggy.

Mandatory Spending

Mandatory spending happens automatically under existing law, without annual appropriations from Congress. Social Security, Medicare, Medicaid, and unemployment compensation are the heavyweights. Together, they make up well over half of all federal outlays. When you see a headline that the budget grows by X percent, most of that growth typically comes from mandatory programs—demographics, healthcare costs, and automatic inflation adjustments push these numbers up regardless of who occupies the White House.

In the budget text, mandatory programs are often called “direct spending.” Look for tables that break out mandatory spending by function, and pay attention to the footnotes. These programs follow their own legislative rules, and the budget proposal may include changes to eligibility or benefit formulas. Those changes aren’t automatic; they need Congress to pass separate legislation, often through a process called reconciliation.

Discretionary Spending

Discretionary spending is the slice Congress votes on each year through twelve appropriations bills. Defense eats up roughly half of this category; the rest covers everything from national parks to medical research to foreign aid. The budget proposal sets a topline number for defense and nondefense discretionary spending, and the Appendix volume drills down into specific accounts.

When you’re reading discretionary figures, notice the difference between budget authority and outlays. Budget authority is the legal permission to commit funds; outlays are the actual cash that goes out the door. A multiyear procurement contract, for instance, may show a big burst of budget authority this year but only a trickle of outlays until the goods arrive. This gap trips up a lot of people. If you want to know the immediate impact on the deficit, focus on outlays. If you want to understand future commitments, watch budget authority.

Net Interest

Net interest is the cost of servicing the federal debt. It’s mandatory in the sense that it must be paid to avoid default, but it’s reported separately because it’s a function of past borrowing and current interest rates, not current policy choices. As the debt grows and rates climb, net interest climbs with it. In recent budgets, it’s become one of the fastest-growing line items. If you spot a budget that projects declining interest costs while debt is rising, check the assumed interest rates with extra care.

Understanding the Baseline

Every budget proposal is measured against a baseline—a projection of what spending and revenues would look like if current law stayed unchanged. The baseline isn’t a forecast; it’s a mechanical calculation. The CBO produces its own baseline, and the OMB produces a slightly different one using the administration’s economic assumptions. Differences between the two can tell you a lot about the optimism baked into the proposal.

When the budget claims it cuts the deficit by a certain amount, that claim is relative to the baseline. If the baseline assumes spending will grow by 5 percent and the budget proposes 4 percent growth, the budget will call that a cut, even though spending is still rising. This isn’t dishonest—it’s standard accounting—but it’s easy to misread. Always ask: cut relative to what?

Person reviewing budget document with a highlighter at a desk

Revenue and Tax Expenditures

The revenue side of the budget matters just as much as the spending side, and it carries its own oddities. Tax revenue comes mainly from individual income taxes, payroll taxes, and corporate income taxes. The budget shows these as receipts. It also includes a chapter, usually in the Analytical Perspectives volume, on “tax expenditures”—revenue losses from special deductions, credits, exclusions, and preferential rates. The mortgage interest deduction, for example, is a tax expenditure.

Tax expenditures work a lot like spending programs, except they’re run through the tax code. A budget that proposes cutting a spending program while expanding a tax credit may just be shifting money from one column to another. Reading the tax expenditure chapter alongside the spending chapters gives you a fuller picture of where government resources are actually directed.

How to Analyze the Tables

The budget is packed with tables. Some are summary tables; others stretch on for dozens of pages. The quickest orientation usually comes from Table S-1 or S-2 in the summary chapter, which shows outlays, receipts, and deficits over a ten-year window. Watch the trend line. Is the deficit projected to shrink, grow, or stay flat? Where are the inflection points? Often, a budget will show deficits shrinking in the first few years and then widening again as assumed economic growth tapers off or temporary provisions expire.

When you move into the detailed account tables, use a simple method: pick a program you care about, find its account number in the Appendix, and trace it back to the summary tables. That shows you how the specific line fits into the broader category and how it’s changed from the prior year. You don’t need to read every line. Sampling a few programs across different functions—defense procurement, Head Start, agricultural subsidies—gives you a feel for the priorities and trade-offs.

Spotting Common Budget Tactics

Experienced budget readers learn to recognize certain recurring techniques. Here are a few to keep an eye out for:

Timing shifts. A payment scheduled for October 1 might get moved to September 30, shifting an outlay from one fiscal year to the next without changing the underlying policy. This can make a single year’s deficit look smaller while leaving the long-term picture untouched.

Emergency designations. The budget may designate certain spending as emergency spending, which is exempt from statutory spending caps. If a proposal labels a recurring need as an emergency, that’s a sign it’s being used to get around budget rules.

Phased-in reforms. A proposal may call for large savings from a program reform, but those savings get pushed into the later years of the ten-year window. The near-term impact is tiny, and the distant savings may never materialize because future Congresses aren’t bound by today’s promises.

Interaction effects. A tax cut can increase the deficit, which raises debt, which raises interest costs. The budget may show the direct cost of the tax cut but not the indirect interest cost. The analytical perspectives volume usually includes a table on “interest effects” that tries to capture these interactions, but it’s worth double-checking.

Following the Money Through Congress

Once the President’s budget is out, the action moves to Capitol Hill. The House and Senate Budget Committees draft a budget resolution, which sets the topline spending and revenue levels. This resolution doesn’t get signed by the President; it’s a concurrent resolution that binds Congress’s own actions. The appropriations committees then fill in the details through twelve bills. Reconciliation instructions, if included, direct specific committees to change mandatory spending or revenue laws to hit the budget targets.

To track what actually becomes law, you need to follow the appropriations process, not the President’s proposal. The CBO publishes scorekeeping reports, and the Government Accountability Office issues reports on specific programs. They’re dry documents, but they reflect enacted law rather than proposed wishes. For anyone serious about understanding federal spending, they’re essential.

Frequently Asked Questions

What is the difference between the deficit and the debt?

The deficit is the annual gap between what the government spends and what it collects in revenue. The debt is the accumulation of all past deficits, minus any surpluses. When the budget shows a $1 trillion deficit, that means the debt will grow by roughly $1 trillion that year, plus or minus certain accounting adjustments. The budget proposal will include a table showing debt held by the public—the portion owed to outside investors—and gross debt, which includes money the government owes to itself, like the Social Security trust fund.

How can I tell if the economic assumptions are reasonable?

Compare the administration’s assumptions to the CBO’s economic projections, which are independent and publicly available. Look specifically at real GDP growth, the unemployment rate, inflation as measured by the Consumer Price Index, and interest rates on 10-year Treasury notes. If the administration assumes significantly faster growth or lower interest rates without a clear policy rationale, the budget numbers may be overly rosy. Also check the date of the assumptions; a budget released in March may rely on economic forecasts from November that have already been overtaken by events.

Why do some programs appear to grow even when the budget claims cuts?

This goes back to the baseline. If current law assumes a program will grow by 10 percent and the budget proposes 6 percent growth, the budget will report a 4 percent cut relative to the baseline—even though actual spending still increases. This is especially common in healthcare programs, where underlying costs rise each year. To see what’s really happening, compare the proposed outlay levels year over year, not just the baseline comparisons.

Where can I find the actual historical numbers?

The Historical Tables volume, released with each budget, provides actual outlays and receipts for past fiscal years, going back to 1940 in many cases. These tables aren’t projections; they’re audited figures. Table 1.1 shows the summary of receipts, outlays, and surpluses or deficits. Table 3.1 shows outlays by function. These are the cleanest numbers in the entire budget document and a good place to ground yourself before diving into the projections.

Building a Reading Routine

You don’t need to read the entire budget to be informed. A disciplined approach works better. On release day, read the President’s message and the summary chapter. Spend an hour with the economic assumptions and the topline tables. Identify three or four programs or policy areas that matter to you and trace them through the Appendix. Then set the documents aside and wait for the CBO’s analysis, which usually appears a few weeks later. The CBO report will re-estimate the President’s proposals using its own economic assumptions and will highlight differences.

The federal budget isn’t a mystery to solve once and for all. It’s a recurring puzzle, released anew each year, with shifting pieces and familiar patterns. Once you learn the structure and the vocabulary, you can read it for what it is: a map of the administration’s intentions, drawn with numbers rather than words. And like any map, it deserves a careful and questioning eye.

Why Data Journalists Need Better Design Standards

Data journalism occupies a tense intersection. On one side, reporters hunt down and analyze complex information. On the other, they have to present it to readers already drowning in numbers. In a lot of newsrooms, design gets treated like an afterthought—a quick coat of paint slapped on once the data is crunched. Jerome Leland argues that this approach undermines the very point of evidence-based reporting. A chart that confuses instead of clarifying means the story fails. Better design standards aren’t a luxury. They’re a core requirement for accurate public understanding.

Close-up of colorful data charts and graphs on a desk with a cup of coffee, representing the visual tools of data journalism

The Visual Gap in Modern Reporting

Walk through any major news website and you’ll find a familiar failure. A deep investigation of municipal spending might include a bar chart with clashing colors and unlabeled axes. An election map might use gradients so subtle that key regional differences vanish. These aren’t just aesthetic flaws. They actively misrepresent the underlying data. A reader who squints at a badly scaled line graph walks away with a distorted sense of trend magnitude. A viewer puzzling over a confusing scatter plot may simply give up—and miss the central finding completely.

The trouble comes from a disconnect between quantitative analysis and visual communication. Journalists who excel at obtaining and cleaning datasets often lack formal training in graphic design or perceptual psychology. Meanwhile, graphic designers assigned to news projects may not fully grasp the statistical shadings they’re supposed to illustrate. The result is a product that satisfies neither side. The data is sound, but the container is cracked.

Take the common practice of using default software templates. A reporter quickly generates a chart in a spreadsheet program and embeds it with almost no adjustment. The colors are factory defaults: a bright blue, a garish orange, a pale green that vanishes on white backgrounds. The legend sits awkwardly in a corner. The font is too small for mobile screens. This shortcut saves minutes in production but costs far more in reader comprehension. When a newsroom treats a chart as a simple export rather than a designed artifact, it underserves the audience.

Perception Shapes Interpretation

Human vision doesn’t passively record what it sees. It actively constructs meaning from shape, color, and contrast. Data visualization researchers have documented this for decades. A study by Cleveland and McGill back in the 1980s showed that people judge the relative sizes of bars more accurately than the relative sizes of circles. And yet news sites keep using bubble charts for precise comparisons. The same body of research shows that color scales must be chosen carefully to avoid introducing false patterns. A traffic light palette—red, yellow, green—can imply a moral judgment where none should exist. A rainbow scale can create artificial boundaries in continuous data.

These perceptual quirks aren’t intuitive. A journalist who hasn’t studied them might pick a color scheme based on personal taste or brand guidelines. That choice can then warp the whole story. For example, a heat map of air pollution that uses a red-to-blue scale might make moderately polluted areas look dangerously severe, simply because red triggers an alarm response. A more appropriate single-hue scale, ranging from light gray to dark orange, would convey intensity without false drama. Design standards have to encode these lessons so that every chart in a newsroom meets a minimum threshold of perceptual accuracy.

Typography also plays a part. Axis labels, annotations, and titles must be legible across devices. A study in the Journal of Usability Studies found that smaller font sizes on mobile graphs significantly reduced data recall. If a reader can’t comfortably read the numbers, they can’t trust the story. Design standards should specify minimum type sizes, contrast ratios, and spacing rules—not as rigid dogma, but as a baseline for clarity.

Person analyzing data visualizations on a wall with sticky notes, illustrating the planning process behind effective data journalism

Why Templates Fail Without Principles

Some news organizations have tried to solve the design problem by building internal charting tools. These systems let reporters drop data into a template and publish a standardized graphic. On the surface, it seems efficient. In practice, it often spawns new headaches. A template that works for a simple time series might be completely wrong for a categorical comparison. A template designed for desktop screens might break on a phone. Without guiding principles behind the tool, the journalist becomes a button-pusher rather than a communicator.

A better approach is to pair templates with a design system that explains why certain choices matter. That system would include rules for picking chart types based on the data structure and the question being asked. It would define color palettes for different variables: sequential for ordered data, diverging for data with a meaningful midpoint, categorical for distinct groups. It would require that every chart include a clear headline stating the takeaway, not just a generic description. A chart titled “Unemployment Rate, 2020–2024” is far weaker than “Jobless Claims Fell Sharply in 2023 but Remained Above Pre-Pandemic Levels.”

Design standards must also address interactivity. When a newsroom adds hover tooltips, filterable views, or animated transitions, it introduces new layers of complication. An interactive map that lets users explore county-level data can be powerful, but only if the controls feel obvious and the feedback is immediate. A study by the Reuters Institute for the Study of Journalism found that readers often abandon interactive features that aren’t self-explanatory. Standards should mandate usability testing before publication, especially for complex projects.

The Cost of Inconsistency

Inconsistent design erodes trust. A reader who sees one chart style in an article about housing prices and a completely different style in a follow-up on rental markets may wonder if the data even comes from the same reliable source. Visual consistency signals editorial rigor. It tells the audience that the newsroom has a coherent method for handling evidence. When a publication uses ten different shades of blue across its charts, or switches between serif and sans-serif labels without reason, it looks sloppy. The content may be excellent, but the presentation undercuts authority.

This isn’t about making every chart look identical. Variety is necessary to match different data types and narrative tones. But that variety should operate inside a defined framework. A newspaper’s style guide governs punctuation, capitalization, and attribution. A data design guide should govern visual grammar with equal care. It should answer questions like: When do we use a line chart versus a dot plot? How do we label outliers? What’s our policy on truncated axes? These aren’t minor details; they’re ethical choices that shape how the public understands numbers.

Consider the truncated axis debate. A bar chart that starts at a value other than zero can exaggerate differences. Some newsrooms ban the practice entirely. Others allow it with a clear visual break and a note to the reader. There’s no universal right answer, but there is a universal need for a clear, documented policy. Without one, individual reporters make ad-hoc decisions that can lead to accusations of bias. A design standard removes ambiguity and protects the newsroom’s credibility.

Team of journalists discussing data charts on a whiteboard, highlighting collaborative design review in a newsroom

Training and Newsroom Culture

Standards on paper mean little if the newsroom doesn’t invest in training. Data journalists need to understand the principles behind the rules, not just memorize them. A one-time workshop won’t cut it. Regular sessions that review both excellent and flawed examples build a shared vocabulary. When a reporter can explain why a particular color ramp distorts a map, they become a stronger advocate for their own work. When an editor can spot a misleading dual-axis chart at a glance, the publication’s quality ticks upward.

This training should extend beyond the data team. Photo editors, copy editors, and social media managers all handle visual content. A social media graphic that crops a chart incorrectly can create viral misinformation. A headline superimposed on a data visualization must not obscure critical labels. Cross-departmental design standards ensure that the integrity of the data survives every platform it touches.

Mentorship is also key. Less experienced journalists often feel pressure to publish quickly and may skip design reviews. A culture that pairs new reporters with visualization mentors can catch errors early. These mentors don’t need to be professional designers. They need to be colleagues who have internalized the newsroom’s standards and can offer constructive feedback. The goal isn’t to slow down the news but to build quality assurance into the workflow.

Accessibility as a Baseline

A design standard that ignores accessibility is incomplete. Millions of readers have visual impairments, color blindness, or cognitive disabilities that affect how they process graphics. A chart that relies only on color to distinguish categories will be illegible to someone with red-green color blindness. A complex infographic without a descriptive text alternative will be invisible to a screen reader user. The Web Content Accessibility Guidelines provide a technical foundation, but newsrooms need to translate those guidelines into practical, journalism-specific rules.

For example, every data visualization should have a companion text summary that conveys the main point. This benefits not only disabled readers but also those who encounter the chart in a low-bandwidth environment or a clipped preview. Patterns and textures can supplement color differences. Interactive elements must be operable by keyboard alone. These aren’t niche concerns. An estimated one in twelve men has some form of color vision deficiency. Designing for inclusivity expands the audience and reflects a public-service mission.

Moving Toward a Professional Standard

The field of data journalism has matured rapidly over the past decade. Investigative teams now routinely use statistical modeling and geospatial analysis. But the visual layer hasn’t kept pace. Many newsrooms still lack a dedicated data designer or a formal review process for charts. This gap isn’t sustainable. As readers become more sophisticated—and more skeptical—they expect the same rigor in presentation that they demand in reporting.

Industry organizations can help. Conferences and awards often celebrate innovative data projects, but they rarely evaluate design consistency across a publication’s entire output. A newsroom that produces one stunning interactive feature per year while publishing dozens of confusing daily charts is not truly serving its audience. Recognition should also go to organizations that maintain high standards across the board, in routine graphics as well as special projects.

Journalism schools bear responsibility, too. Many programs now offer courses in data reporting, but visual design remains an elective. Every aspiring data journalist should graduate with a basic competency in chart selection, color theory, and accessibility. They should know how to critique a visualization and how to iterate on a draft. These skills are as fundamental as fact-checking or interview technique.

In the end, better design standards are about respect. Respect for the data, which deserves an accurate representation. Respect for the reader, who deserves a clear path to understanding. And respect for the craft of journalism, which at its best illuminates the world with precision and honesty. The tools and knowledge already exist. What’s needed is the institutional will to make design a first-class concern, not an afterthought.

Frequently Asked Questions

Why can’t data journalists just use the default chart settings in their software?

Default settings are generic by nature. They aren’t optimized for the specific data or the narrative context of a news story. Colors may be inaccessible, scales may be misleading, and typography may be too small. Journalists have a responsibility to adapt every visual element to serve clarity and accuracy. Relying on defaults outsources editorial judgment to a software engineer who never saw the dataset.

What is the single most impactful change a small newsroom can make?

Start with a short, written style guide for data visuals. It doesn’t need to be long. Include rules for chart types, a limited color palette, minimum font sizes, and a policy on axis scaling. Even a one-page document can dramatically improve consistency. Pair it with a quick review step before any chart goes live—ideally by a second set of eyes.

How do design standards relate to journalistic ethics?

Visual design isn’t neutral. Choices about color, scale, and layout can emphasize or downplay aspects of the data. A poorly designed chart can mislead as effectively as a poorly written sentence. Ethical journalism requires that every element of a story—text, image, and graphic—present information truthfully. Design standards are an extension of the newsroom’s commitment to accuracy and fairness.

Are there legal requirements for accessible data visualizations?

In many jurisdictions, web accessibility is mandated by law for public-facing sites, including news outlets. The Americans with Disabilities Act in the United States and similar legislation elsewhere require reasonable accommodations for users with disabilities. Beyond legal compliance, accessible design aligns with the core journalistic value of serving the entire public. It’s both a practical and an ethical necessity.

The Problem With Truncated Y-Axes in Political Charts

Political charts are inescapable. They flash across cable news segments, fill social media feeds, stuff campaign mailers, and prop up arguments in congressional hearings. The promise is simple: take a messy dataset—unemployment figures, deficit projections, polling swings, GDP growth—and make it legible at a glance. But one quiet design move can flip an honest graphic into a tool of manipulation. That move is truncating the y-axis. The practice, often barely noticeable, reshapes what people think they saw without changing a single number. Knowing why it happens, how to catch it, and what it does to political conversation matters for anyone trying to read the news with their eyes open.

Close-up of a bar chart on a screen with a truncated y-axis

What a Truncated Y-Axis Actually Means

In an honest bar or line chart, the y-axis starts at zero. That baseline gives your eye a fair sense of proportion. When the axis is truncated—when it kicks off at some number well above zero—the visual gaps between data points blow up. A shift from 50 to 55 can be drawn to look like a doubling if the axis begins at 48. The numbers stay correct. The picture lies.

This is not a fresh scandal. Statisticians and data visualization folks have been yelling about truncated axes for decades. Darrell Huff’s 1954 book How to Lie with Statistics gave a whole chapter to the “gee-whiz graph,” where lopping off the bottom of a chart exaggerates tiny wiggles into dramatic swings. Yet the trick hangs on, especially in political messaging, because it lands. People remember the steep line. They forget the axis labels.

Why Political Charts Are Especially Vulnerable

Political communication runs on contrast. A candidate needs to show that crime exploded under an opponent or that their own policies delivered a roaring recovery. Truncating the y-axis supplies that visual jolt without inventing a single number. The chart is, strictly speaking, correct—the figures are real—but the emotional punch is manufactured.

Take a chart of federal spending over four years. Set the y-axis from $3.8 trillion to $4.2 trillion, and a 5% bump looks like a sheer cliff. Start the axis at zero, and that same bump is a mild incline. Same data. Different frame. Political operatives count on the fact that most viewers absorb the shape, not the scale. The truncated chart becomes a tool of selective emphasis, a way to shout without raising your voice.

Newsrooms aren’t immune. Flat lines don’t pull clicks or hold eyeballs during a two-minute segment. A designer might truncate an axis to spotlight a trend the editors consider the real story. The motive may not be sinister, but the effect on public comprehension is identical. When every modest shift gets amplified, the audience loses the ability to separate genuine upheaval from normal wobble.

Person analyzing a political chart on a tablet with a concerned expression

How Truncation Changes Political Narratives

The most immediate result is inflated perceived change. A 2% polling nudge gets drawn as a landslide. A modest budget increase becomes a spending blowout. This visual exaggeration pours fuel on political polarization. When every statistic looks extreme, compromise starts to feel absurd. Why negotiate over a small policy tweak if the chart screams emergency?

Truncation also chews away at institutional trust. A viewer who later learns the real magnitude of a shift may feel played—not just by the politician who waved the chart around, but by the outlet that broadcast it. Over time, that breeds a low-grade cynicism. People stop believing any data graphic, even the honest ones. The chart turns into just another rhetorical prop, and the numbers underneath lose their power to ground a debate.

There’s a secondary effect on policy itself. Decision-makers who work from truncated charts can misdirect resources. A senator staring at a dramatically spiking crime graph might push for draconian sentencing laws, when the actual increase sits inside normal annual variation. A city council looking at a revenue chart that appears to be in freefall might slash essential services, not realizing the drop is barely a rounding error. The visual fib becomes a policy blunder.

Real-World Examples in Political Communication

During the 2020 U.S. election cycle, campaign social media accounts posted charts comparing job growth across administrations. One widely shared graphic showed a steep upward line for the incumbent and a pancake-flat line for the predecessor. The y-axis, however, started at a value that neatly snipped off the recession and early recovery, making the contrast look far harsher than the raw numbers could support. Fact-checkers flagged it, but by then the image had already circulated millions of times.

Government agencies slip up too. A 2017 report from a state health department used a truncated bar chart to display opioid overdose deaths. By starting the y-axis just below the lowest data point instead of at zero, the chart turned a modest year-over-year increase into what looked like a public health catastrophe. The data were alarming enough on their own; the visual distortion added unnecessary panic and chipped away at the department’s credibility with already skeptical legislators.

International examples are easy to find. During the Brexit debates, both sides deployed truncated charts to dramatize economic forecasts. One chart showing projected GDP loss from leaving the EU cut the axis to make a 2% decline look like an economic cliff. The opposing side used a similarly truncated chart to shrink the same projection. Viewers ended up with two contradictory visual truths, each technically accurate but contextually dishonest.

The Psychology Behind the Deception

Human perception of visual data leans hard on proportional reasoning. When we see two bars and one is twice as tall as the other, we infer the underlying values have a 2:1 ratio. Truncation snaps that proportional link. The bars still represent real numbers, but the ratio of their heights no longer matches the ratio of their values. A bar for 50 may stand twice as tall as a bar for 45 if the axis starts at 40, even though the true ratio is a whisper over 1.1:1.

Cognitive scientists call this the “baseline effect.” Our brains treat the visible bottom of the chart as an implicit zero, even when we consciously read the axis labels. Correcting for the missing baseline takes deliberate mental work that most viewers, flicking through news feeds, never apply. The truncation exploits a cognitive shortcut, making the chart’s emotional wallop wildly disproportionate to its mathematical content.

Color and annotation crank up the effect. A truncated chart shaded with a red “danger zone” or sporting an arrow pointing to a “record high” can slip right past rational scrutiny. The viewer swallows the conclusion—things are getting dramatically worse, or dramatically better—without ever engaging the numbers. Political communicators know this cold. They design charts not for analysis but for persuasion.

Person pointing at a chart on a whiteboard during a presentation

When Truncation Might Be Justified

Not every truncated y-axis is a con. In scientific and financial contexts, where tiny variations carry big consequences, a non-zero baseline can uncover patterns invisible on a full-scale chart. A stock price bobbing between $98 and $102 over a week means nothing on a 0–100 axis but plenty on a 95–105 axis. The difference is transparency: the axis break must be clearly marked, and the context has to justify the zoom.

In political charts, that justification is almost never there. The point of a political chart is nearly always comparative: this candidate’s record against that candidate’s, this year’s spending against last year’s. Comparisons demand a shared baseline. Without zero, the viewer can’t accurately judge the size of a difference. A chart showing unemployment sliding from 6% to 4% is honest on a 0–10% axis but misleading on a 3.5–6.5% axis, because the visual drop exaggerates a modest improvement.

Some argue that truncation is fine when the audience is sophisticated and the axis is clearly labeled. But political charts are built for mass consumption. They appear on television for seconds, in tweets scanned between distractions, on flyers glanced at in doorways. The working assumption has to be that the baseline will be taken at face value. If the designer can’t trust the audience to read the axis, the designer shouldn’t truncate it.

How to Spot a Truncated Y-Axis

Building a habit of checking the y-axis takes only a moment. Before you absorb the message of any bar or line chart, look at the vertical scale. Does it start at zero? If not, ask why. A non-zero baseline isn’t automatically a lie, but it’s a red flag. Next, mentally stretch those bars down to zero. Picture what the chart would look like with a full axis. If the visual impact flips dramatically, the chart is misleading.

Watch for charts that use a zigzag or break symbol on the y-axis. That’s a common signal the axis has been snipped. But plenty of designers skip even that small warning, leaving the viewer with no cue at all. Also be wary of charts that show only the tops of bars, a trick sometimes used to hide the baseline entirely. If you can’t see where the bars begin, you can’t trust the proportions.

Another red flag is a chart that compares two groups but uses different y-axis scales for each. This is a related but distinct sleight of hand. A dual-axis chart can make two unrelated trends appear to move in lockstep, or make one trend look far jumpier than the other. Always check both axes and ask whether the comparison actually holds water.

How to Create Honest Political Charts

For anyone producing political content, the fix is straightforward: start bar charts at zero. If small differences matter, use a different format. A table of numbers, a percentage-change callout, or a line chart with a clearly stated baseline can convey nuance without distortion. If truncation is unavoidable, label it prominently and think about including an inset chart with the full axis for reference.

Designers should also resist the lure of 3D effects, which further warp perception by tilting bars away from the axis. A 3D bar chart layers the baseline problem with perspective distortion, making accurate comparison nearly impossible. Flat, clean, zero-baseline charts are the gold standard for honest political communication.

News organizations can adopt internal standards. A policy requiring all charts that compare quantities to start at zero—with exceptions only for indexed data or rates of change—would head off many unintentional distortions. Editors trained in data visualization ethics can catch problematic charts before they go live. Being upfront with the audience about why a particular axis range was chosen builds trust over time.

FAQ

Why do political charts so often use truncated y-axes?

Political charts are built to persuade, not just to inform. Truncating the y-axis makes differences look larger, which can puff up a candidate’s achievements or make an opponent’s failures look catastrophic. It’s a visual shortcut that exploits how our brains process proportions, and it works even when the axis labels are technically accurate.

Is a truncated y-axis always deceptive?

Not always, but in political settings it usually is. In scientific or financial charts, a non-zero baseline can reveal small-scale patterns that matter. The difference comes down to intent and audience. Political charts are meant for quick consumption by a general audience, where the visual impression carries more weight than the precise numbers. If the baseline isn’t zero, the chart is probably exaggerating the trend.

How can I quickly check if a chart is misleading?

Look at the y-axis. If it doesn’t start at zero, imagine the bars or lines extending down to a true zero baseline. If the visual story changes sharply—a steep slope flattens out, or a huge gap shrinks to nearly nothing—the chart is misleading. Also check for missing axis labels, broken axes, or dual-axis charts with mismatched scales.

What should I do if I see a truncated chart in the news?

First, check the source. Reputable outlets sometimes make mistakes, but they also run corrections. If the chart comes from a partisan source, assume the truncation is intentional. Share your concern with the outlet if you can, and when discussing the data with others, point out the axis manipulation. Visual literacy is a branch of media literacy, and calling out bad charts helps raise the bar for everyone.