Rehearsal Makes Ready

Dense, shifting cloud cover filling the sky, with the roofs of several houses visible along the bottom edge of the frame.
The Changing Winds and Clouds

Case of the Week, Min Wu, PhD · ai-public-health.com

Wisconsin had tornadoes this summer, and some Wisconsinites lost their homes. For two days before the worst of it, I kept checking the forecast — the radar, the watches, the timing. Where is it now? What does the model say it will be doing in six hours? In twelve? I could not do anything to stop the storm. What I could do was watch the simulation running ahead of it.

The Forecast You Can't Trial

Before you plan an outdoor event, you check the weather. What you're actually looking at is a simulation: a computational model of atmospheric physics, run on a supercomputer, projecting temperature, precipitation, and wind hours or days before they arrive. Nobody schedules a "trial Saturday" to see whether it rains. The forecast is the only rehearsal you get.

That rehearsal is newer than it feels. In 1922, a meteorologist named Lewis Fry Richardson tried to calculate a six-hour forecast by hand, using physics equations not so different from the ones running today. It took him roughly six weeks, and the answer was wrong badly enough that nobody attempted the method again for a generation. The first forecast that actually worked came in 1950, when a team running the ENIAC computer produced a 24-hour prediction — and getting there took the machine most of a day, running around the clock through several breakdowns. For decades, the rehearsal and the storm arrived at almost the same time. Simulation didn't become useful because the physics improved. It became useful once the computation could finally outrun the weather it was describing.

That's such an ordinary fact about weather that it's easy to miss how unusual it is as a design principle. Most fields don't get a rehearsal. They get a pilot program, a soft launch, a beta group — versions of "trial Saturday" where the first real exposure is the test. Public health AI, more than most fields, can't afford that. The population on the other end of a badly calibrated risk score or a mistimed alert isn't a beta group. It's whoever the system reaches first.

Where This Comes From

I wrote a short reflection several months ago arguing that new conceptual frameworks in public health often stall — not because the ideas are weak, but because the structures set up to evaluate them are built to minimize uncertainty rather than explore it. Middle-range theories — the bounded, mechanism-level kind, not the grand unifying kind — are exactly the ones simulation can stress-test before they ever touch a real population. Digital twins, agent-based models, synthetic scenarios: none of these replace empirical validation. What they do is move the first real failure earlier, and somewhere safer.

I built PAPO around that premise from the outset — simulate the policy-aware personalization logic against synthetic scenarios before recommending it touch a live population, not after. The documentation for how those simulations get specified and reviewed — what I've been calling CMDDS — has been sitting on GitHub for months now, openly, for anyone doing similar work to use or critique. I won't re-explain the structure here; the repository does that job better than a blog post would. What I want to talk about instead is why the habit behind it — simulate before you deploy, and be able to show your work — just stopped being optional.

"The deep-dive portion of this weekly case is reserved for members. Subscribe for free to unlock the full analysis instantly."