How this forecast is built

Written out in full, because a forecast you can't interrogate is just a guess with a nice font.

Eight sources, not one

Every load fetches the National Weather Service gridded forecast, the Open-Meteo blended forecast, MET Norway, and the raw ECMWF, ICON and GFS model runs. Two more join them where they apply: HRRR, the 3 km American short-range model that is re-run every hour and is the only source fine enough to tell rain at 3pm from rain at 6pm — it only covers North America and only reaches about 18 to 48 hours ahead — and AIFS, ECMWF's AI-based model, which votes quietly until the accuracy log says it has earned more. Each is listed at the bottom of the forecast with a link, so any number here can be checked at its origin.

A weighted median, not an average

For each hour the sources are combined with a weighted median. A median ignores a single wild outlier instead of letting it drag the number, which an average cannot do. The weights change with lead time: the NWS forecast — where a human forecaster has reviewed the models — counts most in the first two days, and ECMWF counts most in the second week, which is where it performs best. Rain chance is handled separately, and answers one narrow question: what are the odds that at least 0.01 inch of rain — the standard threshold for measurable precipitation — lands on your exact spot during the hour. Only estimators of that same event are combined, and how many models happen to agree is not one of them; agreement goes to confidence instead.

Related sources share one vote

Six feeds are not six independent opinions. They are grouped by lineage — ECMWF and MET Norway, which is driven by the ECMWF run outside the Nordics; GFS; ICON; the human-edited weather.gov grid — and the Open-Meteo blended product, which is built from those same models. Inside a lineage the members split a single lineage-sized vote, and the blended product's weight shrinks as more of its constituents show up on their own. Agreement is measured the same way: each lineage is reduced to its own middle value first, so the uncertainty band cannot narrow just because a related feed was added.

Ensembles: how sure is the model of itself?

Comparing sources answers one question — do separate forecasters agree? It cannot answer the other one: is any of them actually sure? For that, a model is run dozens of times from slightly different starting conditions, and the results are compared. This app reads 31 GEFS members and 51 ECMWF members, so the rain chance you see can be a real count of members rather than a proxy, and the temperature range is the range the runs actually produced.

The member count is also a vote on the chance of rain itself, not only on confidence. The share of members that go wet is the only natively probabilistic rain number in the line-up, so it joins the published chances when the headline chance is worked out — the two ensembles voting separately, so the 51-member system does not outvote the 31-member one merely for being bigger. The count of wet deterministic runs deliberately does not vote: those runs come from the same model families as everything else, so counting them again would be the same information wearing a different hat. Its weight starts at zero inside the first few hours — at 25 km the ensembles cannot see the shower that the 3 km short-range model and the radar can — then opens up through the first day and peaks a little above a single source's vote from day three onward, where member spread is the best statement of chance available. The accuracy page keeps a control line with that vote removed, so if it is not helping it comes out.

The interesting case is when the two signals disagree. Every source telling the same story while the members spread out means the models share an assumption and the day is genuinely open — a state the old label could not express, and the one where a confident-looking forecast quietly busts. When that happens the app says so outright: forecasts agree, the day does not.

What the confidence label means

The label answers one question: how much could this forecast still change? It comes from how far apart the independent forecast families are over that stretch of hours — the typical spread in temperature, rain chance and wind — capped by how far out the day is and by how many independent families still reach it (not the raw number of feeds).

  • Likely to hold — the sources tell nearly the same story; plan around it.
  • Mostly settled — the picture should hold, numbers may move a little.
  • Details may shift — timing or amounts could still change; check again on the day.
  • Could change a lot — the sources disagree widely; treat it as a rough idea.
  • Rough outlook only — far enough out that only the general pattern is meaningful.

None of these say anything about whether the weather will be bad. "Could change a lot" on a sunny day still means sunny is the best guess.

What the chart shows

Every chart carries a labelled y axis and marks its maximum. The solid coloured line is the combined view; for rain, a dashed line shows the chance on its own 0–100% scale. The shaded band behind the line is the range across all sources for that hour, so widening uncertainty is visible rather than described. "Show each source" draws one line per source for anyone who wants the raw disagreement.

The next hour comes from radar, not models

Inside the next hour, forecast models are slower than your own eyes. So that row is built from radar observations instead: the last four frames are compared to measure how the echo field is drifting, quadrant by quadrant so a sheared storm does not get one averaged vector, and the field is then carried forward along that motion. Coverage is read over a disc around your point that widens with lead time, because a 60-minute position is less certain than a 10-minute one.

Two independent radar feeds are used: a global tile mosaic for presence and motion, and — inside the US — NOAA MRMS reflectivity sampled upstream of your point for magnitude, which is converted to a rain rate with the standard Marshall–Palmer relationship, using snow-appropriate coefficients below 40°F and refusing to report a rate below 34°F. When the two feeds tell different stories, the step is marked, the wetter reading is kept, and confidence drops a tier rather than the conflict being smoothed away. Radar overrides the hourly ensemble for the first hour only; it never edits the rest of the day, and every radar run is logged next to the ensemble and a plain persistence baseline so its skill can be checked later rather than assumed.

Limits worth knowing

  • Beyond the radar hour, rain start times come from hourly data and are approximate to the hour — the app never claims otherwise. Radar amounts are an estimate from reflectivity, not a rain gauge, and outside the US only timing is available.
  • Government alerts are US-only, taken for your exact point from the National Weather Service. Outside the US no alerts will appear, which is not the same as no risk.
  • Air quality is modelled forecast data, not a reading from a monitor near you. The category wording and thresholds are the official US AQI ones.
  • UV is shown as a peak window rather than a burn time, because burn time depends on skin and behaviour this app cannot know.
  • Rain chance is a defensible heuristic, not a calibrated probability. Nothing here has been checked against observations yet, so "60%" means "the sources add up to about 60%", not "it rained on 60% of past days like this". Calibrating that needs a stored record of every issued forecast measured against station reports over months.
  • Saved locations live in this browser only. There is no account and nothing is synced.
Back to the forecast
Add to home screen

Add Ensemble to your home screen: full screen, instant loads, and notifications.

From the address bar:

  1. Click the install icon at the right of the address bar
  2. Confirm “Install”