best weather forecast models per country. table with metrics on different leads
This is an evidence map, not a universal ranking: 35 extracted winner claims plus four methodology sources, covering seven geographies and six forecast-lead bands, current as of 1 September 2026. Raw NWP models, AI models, post-processing systems and consumer providers are scored on different metrics against different observations, so cells are labelled by evidence tier and left blank where no study exists.
The Europe result pools stations across 13 countries — Germany, France, Great Britain, Spain, Italy, Poland, Netherlands, Belgium, Switzerland, Norway, Austria, Czech Republic, Denmark — and is vendor-reported; the underlying station study verifies eleven physical and AI systems over ten months. arxiv.org
The 27 Jua pages behind the Europe claim repeat one finding — EPT-2 beats ECMWF HRES on wind, temperature and solar at 0–240 h — and are folded into a single matrix entry here. jua.ai
Brazil's GraphCast-vs-IFS comparison spans four subregions, three upper-air variables and 6–240 h, but publishes no complete lead-by-lead numeric table, so score cells stay blank rather than inferred. arxiv.org
ForecastWatch's provider study covers 2,100+ locations in eight regions and 25 providers; The Weather Channel's 46.61% global and 42.93% US first-place shares rank providers, not the raw models beneath them. forecastwatch.com
One raw-NWP head-to-head survives: GFS gave 2–3 extra days of lead time over ECMWF for a California atmospheric-river landfall, a single case study rather than a ranking. noaa.gov
MAE and RMSE are error sizes in physical units (°C, m/s) — lower is better. ACC is a pattern-correlation skill score — higher is better. Share of first-place points counts how often a consumer provider ranked first across metrics, regions and lead days — higher is better.
These measure different things against different observations and must never be compared numerically across studies: a 1.06 °C MAE in Czechia says nothing about a 46.61% first-place share globally. Raw NWP output, AI emulators, station-level post-processing and consumer feeds sit at different points of the forecast chain.
| Geography | Variable · metric | Lead | Winner | Type | Evidence tier | Source |
|---|
Evidence map from 35 extracted winner claims, 6 benchmark rows, 8 provider-ranking rows and 4 methodology sources; research current as of 2026-09-01. Cells show the best-supported winner per geography and lead band; a blank cell means no credible evidence was found, not a tie. MAE/RMSE lower is better; ACC and first-place share higher is better. 26 near-duplicate vendor pages and single-station case studies (Kologo GHI, Kentucky Mesonet, northern-tropical-Africa precipitation, Northern Hemisphere upper-air) were cut for space.