Rehearsal Makes Ready
Case of the Week, Min Wu, PhD · ai-public-health.com
Wisconsin had tornadoes this summer, and some Wisconsinites lost their homes. For two days before the worst of it, I kept checking the forecast — the radar, the watches, the timing. Where is it now? What does the model say it will be doing in six hours? In twelve? I could not do anything to stop the storm. What I could do was watch the simulation running ahead of it.
The Forecast You Can't Trial
Before you plan an outdoor event, you check the weather. What you're actually looking at is a simulation: a computational model of atmospheric physics, run on a supercomputer, projecting temperature, precipitation, and wind hours or days before they arrive. Nobody schedules a "trial Saturday" to see whether it rains. The forecast is the only rehearsal you get.
That rehearsal is newer than it feels. In 1922, a meteorologist named Lewis Fry Richardson tried to calculate a six-hour forecast by hand, using physics equations not so different from the ones running today. It took him roughly six weeks, and the answer was wrong badly enough that nobody attempted the method again for a generation. The first forecast that actually worked came in 1950, when a team running the ENIAC computer produced a 24-hour prediction — and getting there took the machine most of a day, running around the clock through several breakdowns. For decades, the rehearsal and the storm arrived at almost the same time. Simulation didn't become useful because the physics improved. It became useful once the computation could finally outrun the weather it was describing.
That's such an ordinary fact about weather that it's easy to miss how unusual it is as a design principle. Most fields don't get a rehearsal. They get a pilot program, a soft launch, a beta group — versions of "trial Saturday" where the first real exposure is the test. Public health AI, more than most fields, can't afford that. The population on the other end of a badly calibrated risk score or a mistimed alert isn't a beta group. It's whoever the system reaches first.
Where This Comes From
I wrote a short reflection several months ago arguing that new conceptual frameworks in public health often stall — not because the ideas are weak, but because the structures set up to evaluate them are built to minimize uncertainty rather than explore it. Middle-range theories — the bounded, mechanism-level kind, not the grand unifying kind — are exactly the ones simulation can stress-test before they ever touch a real population. Digital twins, agent-based models, synthetic scenarios: none of these replace empirical validation. What they do is move the first real failure earlier, and somewhere safer.
I built PAPO around that premise from the outset — simulate the policy-aware personalization logic against synthetic scenarios before recommending it touch a live population, not after. The documentation for how those simulations get specified and reviewed — what I've been calling CMDDS — has been sitting on GitHub for months now, openly, for anyone doing similar work to use or critique. I won't re-explain the structure here; the repository does that job better than a blog post would. What I want to talk about instead is why the habit behind it — simulate before you deploy, and be able to show your work — just stopped being optional.
Audits Just Caught Up
As of April 2026, HHS requires every high-impact AI system used across its divisions — CDC, FDA, NIH included — to show documented proof of bias mitigation, outcome monitoring, and human oversight before it can keep running. Systems that can't produce that documentation get paused or phased out. That's a binding deadline, not an aspiration.
Read plainly, that's an audit requirement. Read as someone who has spent two years building frameworks around simulate-first design, it looks like something else: the paperwork just started asking for exactly what a simulation-first habit was already producing. A synthetic-population stress test. A documented comparison of competing rules. A recorded failure mode. None of that was originally built for an auditor. It was built because deploying an untested mechanism on a real community felt reckless. Reckless and non-compliant, it turns out, are converging on the same checklist.
I don't think that convergence is something to be smug about. I think it's a sign the field is catching up to something that should have been true all along: if you can't show what your system does under a resource-scarce scenario, an equity-stressed scenario, and a plain baseline — three conditions, not one — you don't actually know what it does yet. You've just been lucky so far.
What the Forecast Doesn't Promise
Here's where the weather analogy has to be honest about its own limits. A tornado forecast rests on physics we understand extremely well — fluid dynamics, pressure gradients, decades of validated models. A public health AI simulation rests on something softer: a synthetic population built from assumptions about how people actually behave, assumptions that can be wrong in ways the simulation has no way of flagging to itself. A forecast can miss the exact path by a few miles and still be right about the mechanism. A simulation can run cleanly, pass every gate, and still be wrong about the one variable that mattered — because nobody thought to model it.
That's not an argument against simulation-first. It's an argument against treating it as a finish line. The forecast doesn't stop the storm; it changes what you do before it arrives. A cleared simulation doesn't guarantee a safe deployment; it changes what you're allowed to learn before real people are exposed to it. The outcome monitoring HHS now requires after deployment isn't redundant with the simulation that happened before it. It's the part of the forecast that only becomes available once the weather actually arrives.
Observer's Insight
It never occurred to me, watching the radar during those tornado watches, that I was standing at the far end of a rehearsal that took decades to become trustworthy — Richardson's six weeks by hand, Charney's all-night run on the ENIAC, and everything since, compressed into an app I check without thinking. It also never occurred to me that I was rehearsing the same intellectual move I make every week in this newsletter: trusting a model enough to prepare, but not enough to stop paying attention once it says the coast is clear. Simulation-first isn't a new kind of caution. It's a habit that took most of a century to earn in meteorology — and public health AI is now being asked to build the same rigor in a few years, not a few decades.
Rehearsal makes ready.
As a member of the Springer Nature Author Affiliate Program, I may earn a small commission from purchases made through this affiliated link to my book. Support the author by checking out my textbook, Artificial Intelligence in Public Health: https://tidd.ly/4mH9389