Every data journalist knows the moment. You have a clean dataset, a clear chart, and a finding worth sharing. Then comes the hard part: building a narrative that respects the numbers without turning them into a bedtime story. The temptation to reach for a tool like an AI script generator—such as an AI script generator that fits the project—is real, and growing. These tools ingest a spreadsheet or a statistical summary and output a structured script, complete with scene headings, voiceover cues, and dramatic arcs. They promise speed. They promise structure. What they rarely promise, and what journalists must demand, is fidelity to the data.

This article is not a blanket condemnation of automation. It is a method for evaluating what automation produces. The same rigor we apply to chart design—checking axes, questioning baselines, disclosing uncertainty—must extend to the narrative wrapper. Because a script is not neutral. It selects, sequences, and emphasizes. When that selection is done by a model trained on generic story patterns rather than on the specific dataset in front of you, the result can flatten uncertainty, invent causality, and favor the most dramatic available arc. The question is not whether to use these tools. The question is how to read their output with the skepticism of a data editor.

The Case Study: State-Level Unemployment, Month by Month

To make this concrete, I worked with a real dataset: monthly state-level unemployment rates from the Bureau of Labor Statistics, covering January 2023 through March 2026. The data includes all 50 states plus the District of Columbia. It is seasonally adjusted. It is publicly available. It is exactly the kind of dataset a newsroom might use to build a regional economic story or a comparative feature on labor market recovery.

I then did two things. First, I wrote a short narrative script by hand—the kind a data journalist might draft for a three-minute video segment or an annotated scrolling feature. Second, I fed the same dataset and a simple prompt into a commercially available AI script generator and examined what came back. The specific tool used was the Unsloppy AI script generator, a web-based service that accepts structured data and a topic prompt, then returns a formatted screenplay-style narrative. Its public documentation is limited, and the terms of use do not restrict naming it in an editorial evaluation. The comparison is instructive not because the human version is perfect, but because the differences reveal systematic weaknesses in automated narrative construction.

The Human-Crafted Script

Here is the core of the hand-written script, condensed for print:

SCENE 1: WIDE SHOT — U.S. MAP, JANUARY 2023
VOICEOVER: In January 2023, the national unemployment rate sat at 3.4 percent. But national averages hide as much as they reveal. Nevada was at 5.2. Maryland at 2.1. That spread—3.1 percentage points—is the story we’re going to track.

SCENE 2: LINE CHART — FOUR STATES, 2023–2026
VOICEOVER: We selected four states to follow: Nevada, Maryland, Michigan, and Texas. Not because they’re extreme, but because they represent different labor market structures. Nevada is hospitality-heavy. Maryland has a high federal employment share. Michigan is manufacturing-exposed. Texas is large and energy-adjacent. Watch what happens.

SCENE 3: CLOSE-UP — NEVADA LINE DROPS, MICHIGAN LINE FLAT
VOICEOVER: Nevada’s rate fell from 5.2 to 4.1 over three years. That’s a genuine improvement. Michigan’s rate barely moved—4.1 to 4.0. But here’s what the line doesn’t show: Michigan’s labor force shrank by 1.2 percent over the same period. A stable unemployment rate with a shrinking labor force is not stability. It’s a warning.

SCENE 4: SPLIT SCREEN — TEXAS AND MARYLAND
VOICEOVER: Texas added 900,000 jobs. Its unemployment rate rose slightly—from 3.9 to 4.2—because the labor force grew even faster. That’s a healthy dynamic. Maryland’s rate stayed low, but its labor force participation rate declined. Two low unemployment numbers. Two very different stories.

SCENE 5: FULL MAP — MARCH 2026, WITH ARROWS
VOICEOVER: The national rate in March 2026 was 3.8 percent. The state spread narrowed to 2.4 points. But narrowing can mean convergence or it can mean the high-rate states improved while the low-rate states deteriorated. In this case, it was mostly the former. But you can’t know that from the spread alone.

SCENE 6: TITLE CARD — “WHAT WE DON’T KNOW”
VOICEOVER: These numbers are seasonally adjusted. That means the BLS has removed predictable calendar effects. But seasonal adjustment is a model. It revises. The January 2026 figure you see today may shift by 0.2 points when the annual revision arrives next year. Also, state-level samples are smaller than national samples. The error range on Nevada’s 4.1 is wider than on the U.S. 3.8. We’re showing point estimates. The data is trying to tell you there’s fog around every number. Listen for it.

Notice what this script does. It names its selection criteria. It distinguishes between a falling unemployment rate that signals recovery and a stable rate that masks labor force contraction. It explains seasonal adjustment and revision risk. It ends on uncertainty, not on a tidy conclusion. The narrative arc is present—states change, patterns shift—but the arc is subordinate to the data’s internal logic.

The AI-Generated Script

The automated script, by contrast, opened with a different move:

SCENE 1: DRAMATIC MUSIC — DARKENED MAP, FLASHING RED STATES
VOICEOVER: In the wake of economic turmoil, America’s job market faced a crisis. Some states burned while others thrived. This is the story of a divided nation, told through the numbers that matter most.

The generator had detected variance in the dataset and interpreted it as crisis. It had no access to the fact that a 3.1-percentage-point spread in unemployment during a period of 3.4 percent national unemployment is historically narrow, not wide. It saw difference and reached for the most dramatic available frame: division, turmoil, burning versus thriving.

From there, the automated script selected the five states with the highest and lowest unemployment rates in January 2023 and tracked only those. It ignored the middle. It constructed a narrative of “recovery leaders” and “lagging states” without once mentioning labor force participation, industry mix, or sampling error. It introduced causal language: “Nevada’s tourism-dependent economy dragged it down,” “Texas boomed thanks to energy expansion.” The dataset contains no industry-level variables. The model inferred causality from correlation with geographic stereotypes.

The script closed with a triumphant scene: “By March 2026, the gap had closed. America had healed.” The narrowing spread was presented as resolution. The possibility that low-rate states had simply drifted upward—a less satisfying story—was absent. The uncertainty scene from the human script had no equivalent. There was no mention of seasonal adjustment, revision, or sample size.

Where Automation Fails: A Pattern Recognition Problem

What happened here is not unique to this dataset or this tool. It is a structural tendency of automated narrative generators trained on large corpora of human-written stories. These models learn that a story has a beginning, a middle, and an end. They learn that conflict and resolution are preferred. They learn that extreme values are more narratively useful than central tendencies. They do not learn—because the training data rarely teaches—that the most honest data story sometimes concludes with “we cannot yet say.”

The Authors Guild, in its AI Best Practices for Authors, warns that “AI outputs are generic mashups of pre-existing works ingested during training” and that “when you claim authorship in a work, it means” you are responsible for the thinking behind it. That warning applies with special force to data journalism. A script that flattens uncertainty and invents causality is not just aesthetically generic. It is epistemically damaging. It trains audiences to expect false clarity.

Professional screenwriting, as StudioBinder’s guide on how to write a movie script explains, relies on deliberate structural choices: scene headings that convey geography and time, subheadings that signal shifts without breaking flow, formatting that ensures clarity. These conventions exist to serve the story, not to replace it. An AI script generator can reproduce the format—it can insert INT. and EXT., it can label scenes—but it cannot make the deliberate choices that give those labels meaning. It cannot decide that a scene heading should read “CLOSE-UP — NEVADA LINE DROPS, MICHIGAN LINE FLAT” because that specific juxtaposition carries analytical weight. It will instead generate generic labels: “SCENE 3: THE STRUGGLING STATE” or “SCENE 4: THE SUCCESS STORY.” The format is present. The thinking is absent.

A Checklist for Evaluating Any Data-Driven Script

The solution is not to reject automation. It is to read automated output with the same structured skepticism we bring to a chart with a truncated y-axis or an unlabeled color scale. Below is a checklist designed for newsrooms, educators, and anyone who encounters a script or storyboard that claims to be “generated from data.” Run through these questions before you publish, air, or cite.

  1. Does the script name its selection criteria? If it highlights specific cases (states, sectors, demographic groups), can you explain why those and not others? A script that picks extremes without disclosing the selection rule is building drama on an unstated foundation.
  2. Does the script distinguish between correlation and causation? Look for phrases like “led to,” “drove,” “caused,” “triggered.” If the underlying dataset contains only observational variables, causal language is an inference the script has added. Demand to see the evidence for that inference, or strike the language.
  3. Does the script acknowledge what the dataset cannot show? Every dataset has limits: sample size, measurement error, revision schedules, unmeasured confounders. A script that never says “we don’t know” or “this estimate will revise” is overselling its certainty.
  4. Does the narrative arc match the data’s internal shape, or is it imposed? Not every dataset has a climax. Some trends are flat. Some cycles are incomplete. If the script forces a three-act structure onto data that is essentially a plateau with noise, the structure is lying.
  5. Are the visual cues honest? If the script calls for “flashing red states” or “dramatic music,” ask what in the data justifies that emotional register. A 0.3-percentage-point change in a rate with a 0.2-point error margin does not justify an alarm.
  6. Does the script define its terms? “Unemployment rate” means something specific. “Labor force participation rate” means something else. If the script uses terms without defining them, it assumes a level of literacy the audience may not have—and may exploit that gap for dramatic effect.
  7. Does the script disclose its source and vintage? “BLS data” is not enough. Which survey? Which seasonal adjustment? Which release month? Data vintages change. A script built on preliminary January data may be obsolete by March. The script should tell the audience when the data was pulled and that it may revise.
  8. Does the script include a scene for uncertainty? If not, add one. Even a 15-second title card that says “These numbers have error margins. The state-level estimates are less precise than the national figure. Revisions will come.” That scene is not a buzzkill. It is the difference between journalism and theater.

Why the Narrative Wrapper Matters as Much as the Chart

For years, data journalism has focused—rightly—on visual honesty. We have debated truncated y-axes, misleading color scales, and the sins of the dual-axis chart. But the narrative that surrounds the chart is equally capable of distortion. A perfectly honest line chart, placed inside a script that invents causality and hides uncertainty, becomes a prop in a fiction. The audience remembers the story, not the axis labels.

This is why the checklist above is not optional. It is the narrative equivalent of requiring labeled axes, cited sources, and visible error bars. A newsroom that would never publish a chart without a source line should not publish a script without a methodological disclosure. A journalism school that teaches students to critique visualizations should also teach them to critique the storyboards that those visualizations inhabit.

The AI script generator is not the enemy. It is a tool that accelerates a process—drafting narrative structure—that journalists have always done. The risk is not that the tool exists. The risk is that we accept its output without the same scrutiny we apply to every other stage of data reporting. We would not publish a chart generated by an AI that had never seen the underlying data. We should not publish a script generated by an AI that has seen the data but not understood it.

Toward a Shared Standard for Data Narratives

What would a shared standard look like? Transparency is the starting point. Any published data script—whether human-written, AI-assisted, or fully automated—should carry a brief disclosure: how the narrative was constructed, what tool or process was used, what human editorial review was applied. This is not radical. Film scripts carry credits. Investigative reports carry methodology notes. Data narratives should carry the same.

Training is the next step. Newsrooms that adopt AI script generators should train their staff to use the checklist above, just as they train them to spot a misleading y-axis. The skill of reading a script for epistemic honesty is teachable. It requires no advanced statistics. It requires only the habit of asking: “How does this sentence know what it claims to know?”

And then there is the cultural shift. The most honest data story is sometimes the one that says: “Here is what we see. Here is what we cannot see. Here is what would change our interpretation.” That story does not end with a triumphant chord. It ends with an invitation to keep looking. That is not a failure of narrative. It is the defining virtue of data journalism. The question is whether newsrooms are ready to treat that virtue as a requirement, not an aspiration.

The next time you encounter a script that claims to be “generated from data,” run the checklist. If it fails more than two items, it is not a data story. It is a story that happens to have some numbers in it. The difference matters. In a media environment saturated with automated content, the ability to tell the difference is a skill worth building—one that the checklist above is designed to support.

What Happened When We Fed State Unemployment Data to an AI Script Generator