← BACK TO FEED
catastrophe modellinginsuranceclimate riskhallucinationsdiffusion models

Generative AI is reshaping catastrophe modelling, but physics violations and commercial incentives could undermine the whole thing

by Maximilian Schreiner
Insurers are increasingly using generative AI, particularly diffusion models, to produce more precise catastrophe risk assessments by synthetically generating vast numbers of weather scenarios and sharpening spatial resolution beyond what traditional physics-based models can achieve. However, the technology carries risks, including AI "hallucinations" that can produce physically implausible events. Even where the science improves, there is a commercial tension, as insurers may favour models that produce lower loss estimates to justify writing more business rather than those that most accurately reflect risk.

Catastrophe models have been a fixture of insurance and finance since the 1980s. They divide the Earth into grid cells, apply equations for gravity, friction, and water flow, and estimate the likely damage from earthquakes, hurricanes, and floods. The physics is well understood. The constraint has always been computational: finer resolution costs more, so you pick a level of detail and live with the tradeoffs.

Generative AI is starting to erode that constraint in ways that genuinely matter.

Fathom, which operates as a subsidiary of reinsurer Swiss Re, has built a diffusion model trained on roughly a thousand years of existing climate simulations. That sounds like a lot, but it isn't enough to capture the full distribution of rare, catastrophic events. So the model generates additional synthetic scenarios, eventually producing tens of thousands of years' worth of plausible weather events calibrated to a projected 2030 climate. A second model then sharpens the initially coarse 100x100 kilometre output down to 10x10 kilometres, which is fine enough to actually track precipitation patterns at a useful scale. Fathom's scientific director Oliver Wing puts it simply: the technology has fundamentally changed what's achievable.

Verisk has gone in a slightly different direction. Rather than modelling extreme wind and rainfall as separate sequential events, their generative AI approach handles both simultaneously, capturing how they interact spatially. Research chief Jay Guin says this outperforms traditional machine learning when it comes to spatial precision. Moody's RMS is meanwhile using AI to process satellite imagery after wildfires and hurricanes, extracting loss estimates faster than manual methods allow. For rare tail-risk events where historical data barely exists, the ability to generate plausible synthetic scenarios is particularly valuable.

That word, plausible, is where things get complicated.

Diffusion models can produce scenarios that look physically reasonable but quietly break the rules. A synthetic storm might have the right general shape while violating conservation of energy. A flood event might look credible on a map but couldn't actually happen given local topography. Wing himself is blunt about it: these techniques can generate, in his words, "absolute slop." The physics-based models that underpinned cat modelling for decades had a key advantage in that their outputs were constrained by equations that reflected reality. Generative models don't have that built-in check.

This is a live research problem, not a solved one.

There's a separate issue that has nothing to do with the technology and everything to do with incentives. More accurate models could, in principle, open up markets that major firms have historically ignored. Bangladesh, large parts of Brazil, much of sub-Saharan Africa: these regions have significant climate exposure but low asset values, which made them unattractive to model. Better tools could change that calculation.

But there's a less optimistic reading. Improved models might reveal that losses in already-served markets are higher than current estimates suggest. That would mean larger capital buffers, tighter underwriting, higher premiums, or all three. One modeller quoted in the Financial Times put the dynamic plainly: insurers tend to buy whichever model produces the lower loss estimate, because that's the one that lets them write more business. Underwriters want volume. A model that tells you risk is worse than you thought is commercially inconvenient, regardless of whether it's scientifically correct.

Swiss Re estimates that natural disasters caused $220 billion in damage last year. Insured losses accounted for roughly $107 billion of that. The gap between total damage and covered losses is sometimes called the protection gap, and it's been stubbornly persistent for decades. Better models won't close it on their own, particularly if the commercial logic around model selection continues to reward optimism over accuracy.

READ NEXT
This Cybersecurity Index Tracks Real Breaches and Refuses to Invent a Grand TotalCapital One Releases AI Vulnerability Hunter to the Public23 Million Paidwork Users' Data Dumped Online After Alleged March Breach