Five days before Hurricane Melissa made landfall, an AI model said with 80 percent confidence that it would hit Jamaica as a Category 5. Not Haiti. Not as a weak system. Jamaica, at the top of the scale.
That call, made in October 2025 while conventional models were still arguing about whether the Caribbean storm would strengthen at all, is the clearest evidence yet for what Google DeepMind and Google Research have built. The model is called WeatherNext, and the researchers behind it can’t fully explain why it works.
Melissa was catastrophic anyway. Flooding, landslides, a wrecked Jamaica. But forecasters got warnings out earlier to the communities in the storm’s path, which is the entire point of a forecast.
A day of lead time used to cost a decade
In a paper published Thursday in Nature, the researchers report that WeatherNext predicts cyclones with accuracy that outruns existing models by roughly a day. Its three-day forecast is about as good as what older models managed at two days.
Historically, buying that extra day took about 10 years of work, the researchers said.
“Even a few hours can make a difference,” said Mike Brennan, director of the US National Hurricane Center. Evacuations, staging supplies, moving response resources: all of it runs on a clock, and a wrong call is expensive.
“Time is really golden when it comes to those types of decisions, so the ability to push forecast accuracy out as much as a day beyond what we’ve previously been able to do is really valuable,” Brennan said.
The intensity problem nobody had solved
Track and intensity are two different forecasting problems, and AI had only cracked one of them.
Where a storm is going depends on planetary-scale weather: cold fronts, prevailing winds, the whole global picture. How strong it gets depends on local atmospheric and ocean conditions at a much finer scale, said Kate Musgrave, tropical cyclone group lead at the Cooperative Institute for Research in the Atmosphere and an author on the paper.
“That’s something we just don’t get from these global models,” Musgrave said. Earlier AI models handled track well, but “intensity they could not do well at all.”
Both matter, because the gap between a nuisance storm and a major hurricane is intensity. Melissa was the first time the National Hurricane Centre called a Category 5 while the system was still a Category 1.
Training on weather to predict cyclones
Machine learning wants lots of examples. Extreme events, by definition, don’t provide them.
“We don’t have that much cyclone data, but we have a lot of weather data,” said Ferran Alet, a research scientist at Google DeepMind and one of the paper’s lead authors. “So what we did was train a model to be both good at weather as well as cyclones.”
The team ran it against retrospective data before letting it near a live forecast. “The results were so good that we were skeptical that we would actually see that in the real-time demonstration,” Musgrave said. Then forecasters started using it operationally and the numbers held. “I think everybody was surprised at just how well it did,” Musgrave said.
Coarse data, better answers, no explanation
Here’s the part that unsettled people: WeatherNext uses much lower-resolution atmospheric data than traditional models need for intensity forecasting, and it still wins.
“When we told the community that our model was only using relatively coarse resolution, they were shocked, because that means that the lower-resolution inputs capture more signal about what’s going to happen than previously believed,” Alet said.
The model is finding something in that coarse data. What, exactly, nobody knows. “It’s a black box at the end of the day, but that gives physicists a signal that something is happening that was not previously understood,” Alet said.
From 50 scenarios to 1,000
WeatherNext doesn’t emit a single answer. It generates a spread of possible futures for a developing storm, which is how you catch a butterfly effect, Alet said, where a small early deviation compounds into something very different days later.
Last year the model produced 50 scenarios per storm. Now it produces 1,000.
“That’s something that, with our computing power, we simply can’t do with our existing numerical models,” Musgrave said.
One tool, not the tool
Brennan is careful about how much weight to put on any single model, including this one. “There’s no guarantee that one model, because it did well last year or really did well for this particular storm, is necessarily going to be the best model for the next season or the next storm,” he said.
And the humans don’t come out of the loop. “A hurricane is not just a track or an intensity forecast,” Brennan said. “It requires experts to translate that into what the impacts are going to be, and it’s the impacts that kill people.”
Google DeepMind is open-sourcing the WeatherNext models used during hurricane season, so other researchers can run them and build on them. Alet’s hope is that outside eyes find the physics hiding in the black box.
“I’m very excited about scientific discovery,” he said. “I think AI is giving us new tools to poke into the laws of the universe.”