Announcing the AI Model Intercomparison Project (AIMIP), a community effort to support open evaluation of AI climate models. 🌎 It brings together a shared benchmark experiment & dataset to make it easier to compare models side by side over multi-decade simulations. 🧵

We need transparent ways to evaluate how AI climate models perform on long-horizon forecasting. Weather models already have common evals like WeatherBench; AIMIP is a shared benchmark for AI climate modeling in the spirit of the Coupled Model Intercomparison Project (CMIP).
For AIMIP, models forecast the global atmosphere over 1979–2024, using historical data from 1979–2014 for training and leaving the final decade held out for testing. The benchmark focuses on the atmosphere alone, and leaves model architecture choices up to each submitter.
AIMIP evaluates model performance on: ◙ Overall climate averages ◙ Long-term trends ◙ El Niño-related atmospheric responses ◙ Day-to-day variability ◙ Out-of-sample behavior under warmer sea surface temperatures
AIMIP gives AI climate models a common held-out protocol. Early submissions from six groups often beat a physics-based model on historical averages but diverge on long-term warming, a failure mode that average-case scores hide.
Checking sign-in…
Loading comments…