Essay

Weather Apps That Show Forecast Accuracy: What a Public Verification Record Looks Like, and How to Read the Scores

Somewhere on a national weather service website there is a page almost nobody visits, and it lists in plain tables exactly how wrong that agency's forecasts have been, updated continuously, going back years.

Your weather app has the same data and does not publish it graded forecast accuracy scores . Every consumer app stores its own forecast history, and scoring that history against what actually happened is a weekend of engineering rather than a research program.

A verification record is the artifact that closes the gap. It is a dated, public comparison between what an app forecast and what the sky delivered, scored with the same metrics meteorologists have used for half a century.

This piece is a reading guide for that document. By the end you should be able to open any published accuracy page, go line by line, and say whether the numbers on it mean anything.

A forecast verification record is a published scorecard comparing every forecast an app issued against what actually happened. It reports hit rate, mean absolute error, bias, and skill, broken out by lead time.

What A Published Verification Record Contains

A verification record is a table before it is anything else. One row per variable, per lead time, per location, per scoring period — with the sample size printed beside every score.

The variables are the things the app actually promised you: high temperature, low temperature, whether it rained, how much, wind speed, and for a photography-minded app, whether the sunset it called delivered. Each needs its own row, because an app can be excellent at temperature and hopeless at rain timing.

Lead time gets its own axis for the same reason. A day-one forecast and a day-seven forecast are different products, and a record that averages them into one figure has destroyed the most useful thing it had.

Keep in mind that the scoring period matters as much as the score. A record covering a full year, transitional weeks included, tells you something that a record covering one calm August does not.

The Numbers That Carry A Scorecard

Five metrics do most of the work on temperature, and precipitation adds three more of its own. Read them together, because any one of them alone can be gamed or misread.

What Is A Hit Rate?

Hit rate is the share of forecasts that landed inside a stated tolerance. For a high temperature that tolerance might be 2°F; for rain it is usually a yes-or-no call scored against whether measurable precipitation fell at a specific station.

Hit rate is the share of forecasts that landed inside a stated tolerance, such as within 2°F or the correct rain call. It is only meaningful when the tolerance and the sample size are printed beside it.

The tolerance is the whole game. A 92% hit rate within 5°F and a 68% hit rate within 2°F can describe the identical set of forecasts, so a record that omits its tolerance has told you nothing at all.

What Is Mean Absolute Error?

Mean absolute error, written MAE, is the average size of the miss with the direction stripped out. A three-degree miss high and a three-degree miss low both count as three.

Mean absolute error is the average size of the miss, ignoring direction. A 2.4°F MAE means the typical high-temperature forecast was off by 2.4 degrees in one direction or the other across the sample.

MAE is the hardest number on the page to flatter, which is exactly why it belongs there. What's more, it is directly usable: if the record says day-two highs carry an MAE near 3°F, then a forecast of 74 is functionally a forecast of 71 to 77.

You will sometimes see RMSE, or root mean square error, which squares each miss before averaging. RMSE punishes large busts far harder, so a low MAE sitting next to a much higher RMSE tells you the app is usually close and occasionally spectacularly wrong.

What Is Forecast Bias?

Bias is the average signed error, so it keeps the direction that MAE throws away. Add every miss with its plus or minus intact, divide by the count, and you learn whether the app leans warm or cold, wet or dry.

Bias is the average signed error, so it shows direction rather than size. A +1.8°F bias means the app runs warm, and you should read about two degrees off every forecast it hands you.

For a reader, bias is the most immediately actionable number on the page, because you can correct for it yourself. An app with a persistent cold bias in winter is still perfectly usable, provided you know to add the offset before you choose the coat.

Of course, a bias near zero is not automatically a compliment. An app that misses 6°F warm half the time and 6°F cold the other half posts a beautiful zero bias alongside a miserable MAE, which is precisely why the pair has to be read together.

What Is Skill Versus Climatology?

Climatology is the thirty-year normal for that date and place — the number you would guess with no forecast at all. Persistence, meaning tomorrow will resemble today, is the other standard baseline.

Skill versus climatology measures the forecast against the thirty-year normal for that date. Positive skill means the forecast beat that baseline, and a score near zero means it added nothing.

A skill score of 1.0 is a perfect forecast, 0 matches the baseline, and a negative number means the app did worse than guessing the average. This is the metric that separates forecasting from a well-designed number generator.

Skill is also where seasonal honesty shows up. Beating climatology in a San Diego July is close to trivial, while beating it during an April frontal passage in Chicago is the actual test — the kind of setup we unpack in how the jet stream steers your week.

How Much Does Accuracy Decay By Lead Time?

Every metric above degrades as the forecast reaches further out, and the shape of that decay is the most informative curve on the page. The atmosphere is chaotic, so small errors in the initial analysis compound with every hour the model simulates forward.

Accuracy decays with lead time, and an honest record shows the decay curve. Day-one highs may land inside 2°F while day-seven highs drift several degrees, so one app can be reliable at one range and not another.

A record that publishes one blended accuracy figure across all lead times is hiding this curve. Averaging day one with day ten flatters the far end and penalizes the near end, which is convenient for the publisher and useless for you.

For why the models diverge as the range grows, we walk through the major global systems in how forecast models actually work, and the spread between them in what forecast confidence really measures.

What About Rain — POD, FAR, And Reliability?

Precipitation needs its own metrics, because a yes-or-no event is scored differently from a continuous one. Probability of detection, or POD, is the share of actual rain events the app called; false alarm ratio, or FAR, is the share of its rain calls that produced nothing.

An app can drive POD toward 1.0 by forecasting rain constantly, which drags FAR up alongside it, so the two are always read as a pair. Reliability is the third piece: across every day the app said 30%, rain should have fallen on close to 30% of them.

That last property is what separates a calibrated probability from a decorative one, and it is the subject of what a 40% chance of rain actually means.

Reading The Metrics Side By Side

Here is how the core metrics compare in what they answer and where each one can mislead. Read across the row rather than down the column.

MetricWhat it answersAn example of a strong resultHow it can mislead
Hit rateHow often the forecast landed close enoughDay-1 high inside 2°F on most daysTolerance widened until nearly everything counts
Mean absolute errorHow big the typical miss isLow single digits °F at day oneHides direction; one enormous bust barely moves it
BiasWhich way the app leansWithin a few tenths of a degree of zeroWarm and cold misses cancel into a flattering zero
Skill scoreWhether the forecast beat the baselineClearly positive against climatologyEasy climates and calm seasons inflate it
Sample size (n)How much weight the other four can carryHundreds of scored forecasts across a full yearA tiny n makes noise look like performance

All of these add up to a single habit. No one cell in that table is a verdict, and the temperature metrics only mean something when you can see them at once.

How To Read A Scorecard Line By Line

A usable record should look boring. Here is a sample line in the format a real scorecard uses:

Sample row. High temperature, day 1, Chicago O'Hare, June — hit rate 82% within 2°F, MAE 1.6°F, bias +0.4°F, skill 0.61 versus climatology, n = 30.

Read it in order. The variable and lead time say what was promised, the station and month say under what conditions, and n = 30 warns you at the outset that a single month is a small sample.

The bias of +0.4°F says the app ran barely warm, which is not worth correcting for. The MAE of 1.6°F says the typical miss was under two degrees, and an 82% hit rate inside a 2°F tolerance is consistent with that.

The skill of 0.61 is the line that matters most. It says the forecast removed roughly sixty percent of the error you would have carried by simply guessing the June normal, which is a real forecast doing real work.

Now change one number and watch the reading change. If the same row read bias +2.9°F against an MAE of 3.1°F, you would be looking at an app that is almost always warm — irritating, but correctable, since subtracting three degrees gets you a usable answer.

Then change a different number instead. If the bias held at +0.4°F while the MAE climbed to 4.8°F, the app has no lean you can correct for and simply misses hard in both directions, which is the harder failure to live with.

The Four Questions That Decide Whether A Record Is Honest

A scorecard can be arithmetically flawless and still be built to flatter. Before you believe one, put it through these four questions:

  • What is the truth source? Forecasts should be scored against a named observing station or an official gridded analysis. An app that scores itself against its own later forecast is grading its own homework.
  • What is the sample size? Hundreds of scored forecasts across a full seasonal cycle mean something, and thirty days in a settled month do not. Note that n belongs on every row rather than in a footnote.
  • What is the scoring window? A record covering one hand-picked quarter invites cherry-picking, while a rolling twelve-month window that updates automatically does not. Be aware that a static page with no update date is a warning on its own.
  • Is every forecast counted? The record has to include the busts, the days the app hedged, and the days the data feed failed. Any exclusion rule should be stated in plain language, because silent exclusions are where accuracy claims go to hide.

An honest record names its truth source, prints the sample size on every row, uses a rolling window, and counts every forecast including the busts. Missing any of the four makes the score unverifiable.

All of these come down to one principle. A verification record earns trust by being inconvenient for the publisher, and if reading one never makes the app look bad, treat it as a brochure.

Why Almost No Weather App Publishes One

The obstacle is commercial rather than technical. Publishing a record converts a vague impression into a fixed number, and "the most accurate forecast" is a claim that only survives while nobody can check it.

There is a second reason, and it is more sympathetic. Most apps are thin clients over a licensed data feed, so the accuracy they would be publishing belongs to a provider upstream that they do not control.

That said, the reader does not care whose model missed. In fact, "our vendor was wrong" is exactly the accountability gap a verification record exists to close.

National agencies have no such escape, which is why their verification pages remain the best free examples of the form. Both NOAA's National Weather Service and the ECMWF publish ongoing verification for their own products, and an afternoon with either one will show you what an unflattering, honest record looks like.

What Vesper Publishes

We publish the call before the day and the result after it. Sunset Verify scores whether the sky we predicted actually arrived, and the brief archive keeps the original wording rather than quietly editing yesterday's take.

We think most weather apps have this backwards, and we think the record supports us. An app that never revisits its own forecasts is optimizing for the moment you check it rather than for whether you dressed correctly or drove to the overlook for nothing.

Cloud cover is where this matters most, because it is among the least skillful fields any model produces. A five-day cloud icon is close to decorative, which is why the light call gets verified instead of asserted.

The editorial method behind all of it — what goes into a brief, and more importantly what gets cut — is written down in how we write a daily brief.

What To Do With The Score

A verification record is only worth reading if it changes a decision. Here is the decision rule we would hand a reader holding one:

  • Correct for the bias first. If the record shows a persistent warm or cold lean, apply the offset in your head before you pick the layer. A +2°F bias in a shoulder-season city is the difference between a trench coat morning and an overdressed one.
  • Trust the near term as numbers, the far term as ranges. Use day one and day two literally, and read anything past day five as a distribution with the published MAE as its half-width.
  • Judge the variable, not the app. An app with strong temperature scores and weak precipitation-timing scores is still worth keeping; you simply stop asking it when the rain starts.
  • Watch skill seasonally. If skill collapses during your local storm season, that is the season to lean on live weather radar rather than the seven-day.
  • Re-check after any redesign. Providers change blends and post-processing without announcing it, so a record that was true last spring is a hypothesis this spring.

Overall, a published record moves the argument from taste to arithmetic. You stop asking which app feels right and start asking which one has agreed, in writing, to be measured.

For the rest of this thread — including how to grade a forecast yourself when no published record exists, in how accurate is your weather app, and how to turn a forecast into a packed bag in packing by forecast — start at the journal.

Common Questions

Which weather apps publish a forecast accuracy record?

Very few. Most consumer apps publish nothing at all, a small number of independent services score the major forecast providers, and national weather agencies publish verification for their own products. Vesper publishes its own record, including whether each sunset call we made actually delivered.

What counts as a good mean absolute error for temperature?

It depends entirely on lead time and climate. Next-day high temperature errors in the low single digits Fahrenheit are typical of a good forecast, while day-seven errors several times that size are normal rather than a failure. Judge the number against the lead time printed beside it, never on its own.

Why does forecast accuracy get worse further out?

The atmosphere is chaotic, so tiny errors in the starting analysis grow with every hour a model simulates. Forecasts stay sharp for a few days, then converge toward climatology as the signal fades. That is why a seven-day high is best read as a range with the published error as its half-width.

Can an app post a high hit rate and still be inaccurate?

Yes, and it is the easiest number on the page to inflate. Widening the tolerance, scoring only a settled month, or excluding the days the data feed failed all raise a hit rate without improving a single forecast. Always read it alongside the tolerance, the sample size, and the scoring window.

What does skill versus climatology mean in a forecast score?

Climatology is the thirty-year normal for that date and place, the guess you would make with no forecast at all. Skill measures how much of that baseline error the forecast removed. A score near zero means the app told you nothing you could not have looked up, and a negative score means it hurt.

K