Do you know which model produced the forecast on your phone this morning, or what time that model last ran? Almost nobody does, because almost no consumer weather app tells you.
The number you saw — 78 and clear, 40% chance of rain by four o'clock — arrived with no provenance whatsoever. It could have come from a global model initialized six hours ago, a rapid-refresh model updated eleven minutes ago, or a proprietary blend of several models that spent the entire run disagreeing with one another.
That missing context is the subject of this piece. We think a forecast without its source model, its run time, its confidence, and its track record is a marketing claim rather than a prediction, and claims should be auditable.
A transparent forecast shows four things: the spread across model members, the source model and its run time, a stated confidence level, and a published verification record. Anything less is a number without provenance.
What Transparency In A Forecast Actually Consists Of
Transparency has become a marketing word in consumer weather, which means it needs a definition specific enough to fail. Here is the standard we hold ourselves to, and the one we think you should hold every other app to:
- Published spread. The forecast shows the range across ensemble members rather than collapsing them into one averaged number. When the members disagree, you see the disagreement.
- Named source and run time. Every number carries the model that produced it and the initialization hour of that run. A temperature from the 12Z ECMWF is a different object than a temperature from the 18Z GFS.
- Stated confidence. Confidence travels alongside the prediction, in the same sentence, rather than being implied by false precision. A high-confidence 71 and a low-confidence 71 ask different things of your day.
- A verification record. Past forecasts are scored against what actually happened, and the score is published where a reader can find it. Accuracy that is never measured is an assertion.
Each of these is cheap to publish and expensive to fake, which is what makes the set useful as a standard rather than a slogan. An app that meets all four has given you enough to argue with it; an app that meets none has given you a number to obey.
Why A Single Averaged Number Hides The Most Important Information
Modern forecasting runs in ensembles. Rather than solving the atmosphere once, a forecast center perturbs the initial conditions slightly and solves it dozens of times, producing a family of plausible futures instead of a single one.
The ensemble mean — the average of all those members — is the number most apps display. That average is genuinely the best single guess in the statistical sense, and it is also, in the most consequential situations, a forecast that no individual member made.
Consider a coastal storm three days out where the ensemble splits into two clusters. Twenty members carry the low far enough offshore that the city gets cold rain, and twenty-five carry it close enough that the city gets eight inches of snow.
The mean of that distribution is roughly two inches of wet snow, an outcome that is not merely unlikely but structurally impossible, because the storm either tracks close or it does not. Averaging a bimodal distribution invents a third scenario and then presents it with more apparent confidence than either of the real ones.
Averaging ensemble members produces a middle outcome that no member actually predicted. When a storm track splits into two clusters, the mean invents a third forecast that matches neither cluster and matches nobody's day.
This is why we publish the shape of the distribution whenever the shape carries the story. Keep in mind that a reader who knows the outcome is bimodal makes a better decision than a reader handed a confident average.
What Model Disagreement Tells You
The major operational models are not interchangeable, and their disagreements are diagnostic rather than embarrassing. Each one carries different physics, different resolution, different data assimilation, and different well-documented biases.
Here is how the models most likely to sit behind your forecast compare, and what each one is actually for:
| Model | Run by | Cadence and resolution | Best used for |
|---|---|---|---|
| ECMWF (HRES) | European Centre for Medium-Range Weather Forecasts | Headline runs at 00Z and 12Z, approximately 9 km | The medium range, days three to ten; generally the strongest verification scores |
| GFS | NOAA / NCEP | 00Z, 06Z, 12Z and 18Z, approximately 13 km | Global coverage out to sixteen days; faster and freely available, more volatile run to run |
| GEFS | NOAA / NCEP | Four runs daily, 31 members | Probability and spread; shows the range a single deterministic GFS run conceals |
| ECMWF ENS | ECMWF | 51 members | The standard reference for medium-range uncertainty |
| HRRR | NOAA | Hourly, 3 km | The next eighteen to forty-eight hours; resolves individual thunderstorms |
| ICON and UKMO | DWD (Germany) and the Met Office (UK) | Multiple runs daily | Independent tiebreakers when the GFS and the ECMWF split |
When those models agree, the atmosphere is in a predictable regime and you can plan around the forecast. When they split — one carrying a trough east, another holding it back — the honest output is a range, and the width of that range is the most useful number on the page.
Model spread is the honest measure of uncertainty. When the GFS and ECMWF ensembles cluster tightly, confidence is high; when members fan out across 200 miles of storm track, the forecast is a range rather than a number.
Spread also has a shape worth reading, which is why we keep how the jet stream sets forecast reliability and what each forecast model is actually good at as standing references. A flat, fast zonal jet produces tight ensembles, while a buckled blocking pattern produces spread you could drive a truck through.
The Run Time Nobody Shows You
Every model output carries an initialization time, written in UTC and marked with a Z. The GFS initializes at 00Z, 06Z, 12Z and 18Z, the ECMWF high-resolution run at 00Z and 12Z, and the HRRR every single hour.
Those runs also take time to complete and distribute, so the forecast reaching your screen is always older than the clock implies. A 12Z global run is typically not fully available for hours afterward, which means an afternoon forecast can rest on a snapshot of the atmosphere taken before you had breakfast.
The GFS runs four times daily at 00Z, 06Z, 12Z and 18Z, the ECMWF high-resolution run twice daily, and the HRRR every hour. A forecast shown without its initialization time may already be six hours stale.
This matters most in convective weather, where the atmosphere reorganizes faster than the global models refresh. If you are deciding whether a line of storms reaches you before sunset, an hourly HRRR run beats a six-hour-old global run every time.
Knowing which one you are looking at is the difference between a decision and a guess. For the same-hour picture, reading weather radar covers what the models cannot yet see.
Confidence Belongs Next To The Prediction
Consumer weather has one confidence expression that everybody sees and almost nobody understands: the chance of rain. The National Weather Service defines probability of precipitation as confidence multiplied by areal coverage, which means a single percentage carries two entirely different meanings.
If forecasters are nearly certain that rain falls but only across 40% of the forecast area, that is 40%. If rain would cover the whole area but forecasters are only 40% sure it materializes at all, that is also 40%.
Those two situations call for opposite decisions, and collapsing them into one figure is the most common act of hidden averaging in consumer weather. We unpack the arithmetic in what a chance of rain actually means.
Probability of precipitation equals confidence multiplied by areal coverage. A 40% chance can mean near-certain rain over 40% of the map, or a four-in-ten shot at rain falling everywhere.
Confidence also decays with lead time in a way worth stating plainly. Forecast skill runs high through roughly days one to three, useful but softening through days five to seven, and by day ten most model output has drifted toward climatology — the long-run average for that date and place.
An app that shows you an exact hourly temperature fourteen days out is presenting noise in the visual language of precision. Remember that precision without skill still reads as authority, which is exactly why it works on us.
What A Verification Record Proves
Accuracy claims are easy to make and nearly impossible to check, which makes the verification record the load-bearing piece of the whole standard. Meteorology already has the tools for this, and they have existed for decades.
The Brier score measures the mean squared error of probabilistic forecasts, rewarding a forecaster who says 90% and is right while punishing one who says 90% and is wrong. A reliability diagram goes further, plotting forecast probability against observed frequency so a reader can see whether the 70% days actually produced rain about 70% of the time.
A calibrated forecaster's 70% days produce rain roughly 70% of the time. Brier scores and reliability diagrams measure that calibration, and publishing them converts an accuracy claim into an auditable record.
Calibration is a different virtue from sharpness, and both matter. A forecaster who says 50% every day is perfectly calibrated in a climate that rains half the time and completely useless, while a forecaster who commits to 5% and 95% and stays calibrated is doing real work.
Publishing both numbers is what separates a service that has measured itself from one that has asserted its accuracy in an app store description. Centers like the European Centre for Medium-Range Weather Forecasts publish their verification scores openly; the apps repackaging that output rarely do.
How Vesper Applies The Standard
We wrote the four-part standard above because we wanted something we could be held to, in public, by readers who know more meteorology than we do. Here is how each piece shows up in the journal:
Spread over averages. When the ensembles disagree, the daily brief says so in words and describes both outcomes rather than splitting the difference into a number nobody's day will match.
Named source and run time. When a call depends on a specific model run, we name the model and the initialization hour inside the brief, so you can check our work against the same output.
Confidence in the sentence. Every forecast judgment carries its confidence in plain language — high, moderate, low — placed next to the prediction rather than buried in a disclaimer.
A published verification record. Sunset Verify scores our sky calls against what the sky actually did, and the record stays visible whether the call was right or wrong.
None of this makes our forecasts better than the models we read, and that is not the claim. The claim is that you can see what we are reading, when we read it, how sure we are, and how often we have been right.
That combination is worth more than a cleaner-looking number with none of it attached. Our editorial process for turning raw model output into a written take is documented in how Vesper writes a daily brief, and it exists for the same reason as everything else on this page.
What To Ask Of Any Weather App
You do not need a meteorology degree to audit the app already on your phone. Ask it four questions, in this order:
- Where did this number come from? If tapping the forecast never surfaces a model name, the app has decided its source is none of your business.
- When was it made? A visible initialization or update time tells you whether you are looking at this hour's atmosphere or this morning's.
- How sure is it? Look for a range, a spread band, or a stated confidence level somewhere other than the fine print.
- How often is it right? Any service confident in its accuracy can publish a verification record, and the absence of one is itself an answer.
Most apps fail at least three of the four, and several fail all of them while advertising accuracy as the headline feature. That gap is why we started publishing this way in the first place.
If you want the version of this that arrives every morning as a written take rather than a dashboard, that is the daily brief in the Vesper journal — one story about the day, the confidence behind it, and what it means for what you wear and when you shoot. Start with reading surface pressure maps if you would rather check our work from the source.
Common Questions
The questions readers send us most often about forecast confidence and model disagreement:
Because they are reading different models, or the same model at different run times. One app may be showing the 06Z GFS while another shows the 12Z ECMWF, and six hours of difference in initialization can move a storm track by fifty miles. Neither app is lying; they are quoting different sources without telling you which.Why do two weather apps show different forecasts for the same hour?
The National Weather Service defines probability of precipitation as confidence multiplied by areal coverage. A 40% chance can mean forecasters are nearly certain rain falls but only over 40% of the forecast area, or that rain would cover the whole area and they are 40% sure it happens at all. Those two situations call for different decisions.What does a 40% chance of rain actually mean?
Skill stays high through days one to three and remains useful but softening through days five to seven. By day ten, most model output has drifted toward climatology, the long-run average for that date and place. An app showing an exact hourly temperature fourteen days out is presenting noise in the visual language of precision.How far out is a weather forecast still useful?
Spread is how far the individual members of an ensemble forecast disagree with one another. Tight spread means the atmosphere sits in a predictable regime and you can plan around the forecast with confidence. Wide spread means the honest answer is a range, and your plan needs a fallback rather than a single number.What is model spread and why should I care about it?
Ask it four questions: what model produced this number, when did that model run, how confident is the forecast, and how often has it been right. An app that surfaces the model name, the initialization time, a spread or confidence band, and a public verification record has given you enough to argue with it. One that surfaces none of those has given you a number to obey.How can I tell whether a weather app is being honest with me?