Have you ever opened a spaghetti plot? If you follow storms closely enough to have gone past the app forecast and into raw model output, you have almost certainly seen one — a map covered in dozens of thin colored lines, some braided tightly together and some wandering off across the ocean alone.
That tangle is the most honest picture of forecast confidence that operational meteorology makes public. Once you can read it, you can answer a better question than whether the forecast is right: how much the atmosphere itself agrees with it.
We have written before about why forecast models disagree with each other and about what forecast confidence actually means when a brief calls a day likely. This is the next layer down — the spread inside a single model run, member by member.
Ensemble spread is the disagreement among members of one model run. Tight, bundled lines on a spaghetti plot mean high confidence, and lines that fan apart mean the atmosphere is unresolved at that lead time.
What An Ensemble Actually Is
A deterministic model run takes one best estimate of the current atmosphere and integrates it forward in time. An ensemble takes that same estimate, perturbs it dozens of times within the known uncertainty of the observing network, and runs every version to completion.
Each of those runs is a member. NOAA's Global Ensemble Forecast System carries 31 members, the ECMWF ensemble carries 51, and the Canadian global ensemble carries 21 — counts that have crept upward as computing has gotten cheaper.
The perturbations are deliberately tiny. A fraction of a degree in a radiosonde temperature over the Pacific, a slightly different sea surface analysis, a stochastic nudge to the model's own convection scheme — at hour zero the members are close to indistinguishable.
By day five, those differences can separate a dry weekend from three inches of rain. That amplification is the whole reason ensembles exist, and it is the working face of what Edward Lorenz found in the early 1960s when a rounded set of initial conditions sent his simulation down an entirely different path.
From One Forecast To A Distribution Of Outcomes
A single run gives you one number for Thursday's high and no way to judge how fragile that number is. An ensemble gives you fifty-one numbers, and the shape of those fifty-one is the actual forecast.
This is where probability language comes from in the first place. When a forecast says there is a 40 percent chance of rain, some version of that number was produced by counting the members that generated measurable precipitation at your location.
Keep in mind that the counting is doing the work, not any single line. Twenty of fifty-one members putting a surface low over your city is a real, roughly 39 percent scenario, even when the headline deterministic run keeps that low two hundred miles offshore.
Count members rather than eyeballing lines. If 40 of the 51 ECMWF ensemble members bring rain to your city, that is roughly a 78 percent chance, and the member count is the probability.
How To Read A Spaghetti Plot
Most spaghetti plots draw one contour per member on a single map. The usual choice is the 500 hPa geopotential height field — commonly the 5,460-meter and 5,640-meter lines — because mid-level heights describe the steering flow that pushes surface systems around.
The trick is reading the plot in passes rather than trying to absorb the whole tangle at once. Here is the order we work through, and what each pass is looking for:
- Bundle width. Measure how far apart the members sit at the feature you care about, not across the whole map. A ridge axis spread over eighty miles and a trough axis spread over six hundred are two different forecasts on the same image.
- Where the bundle loosens. Follow the lines upstream until they tighten again. The point where agreement breaks down is usually the feature driving your uncertainty, and it is frequently a shortwave over the Pacific or a jet stream disturbance nowhere near you.
- Cluster count. Ask whether the fan is one broad smear or two distinct families of solutions. Two families is a fundamentally different situation from one wide one, and it changes what you should do about it.
- Outlier behavior. Note the members sitting well outside the pack and check whether the same members are outliers on consecutive runs. A lone extreme solution that keeps reappearing deserves more attention than one that shows up once and vanishes.
- Change since the last run. Compare today's midday plot with yesterday's at the same valid time. A bundle that is tightening run over run is converging on an answer, and a bundle that is widening is telling you the atmosphere has not decided.
All of these passes serve one question: where, specifically, does the uncertainty live? A plot that looks chaotic at first glance often turns out to be tightly agreed on everything except the timing of one front, which is a forecast you can actually plan around.
Three Kinds Of Spread, Three Different Responses
Disagreement comes in three distinct flavors, and identifying which one is in front of you matters more than measuring how much of it there is. The distinction changes the decision, not merely the confidence level.
Position Spread
Position spread appears when members agree that something happens and disagree about where. A hurricane forecast cone is the most familiar rendering of this, though the cone itself is drawn from historical error statistics while the ensemble underneath it is a set of members arguing about track.
This is the friendliest kind for a reader, because it translates directly into a personal answer. If your city sits at the edge of the member envelope rather than in the middle of it, your risk is genuinely lower than the headline suggests.
Timing Spread
Timing spread appears when members agree on what and where and disagree about when. A frontal passage that arrives at three in the afternoon in half the members and at nine in the evening in the other half produces identical daily totals and completely different afternoons.
For anyone planning around light, this is the kind that hurts. Six hours of disagreement about when clouds clear is the difference between a spectacular golden hour and a flat gray sky, and no daily-summary forecast will ever surface it.
Existence Spread
Existence spread appears when members disagree about whether a feature forms at all. Thirty members develop a coastal low and twenty-one keep the energy sheared apart, and there is no meaningful average of those two worlds.
Existence spread carries the highest stakes and produces the least useful mean. When you see it, the honest forecast is a pair of named scenarios with probabilities attached rather than a single number.
What The Ensemble Mean Hides
The ensemble mean is the most widely republished ensemble product and the easiest one to misread. It reports the center of the distribution and says nothing whatsoever about that distribution's shape.
Consider a coastal storm in which twenty-five members carry the low up the shoreline and twenty-six push it out to sea. The mean shows a weak, smeared system splitting the difference — an outcome that not one of the fifty-one members forecast.
The ensemble mean reports the center of the distribution and nothing about its shape. When members split into two clusters, the mean lands between them and describes an outcome no individual member predicts.
Averaging also flattens gradients everywhere it is applied. Precipitation maxima get smeared across the spread of member positions, which is why ensemble mean rainfall totals almost always undersell the peak that the eventual winning solution produces.
Be aware that this cuts the other way as well. Ensemble mean fields are smooth and reliable for large-scale pattern recognition, and we use them constantly for exactly that — the failure appears only when someone reads the mean as a point forecast.
Turning Spread Into A Decision
Reading the plot is half the work, and having a rule for what each picture means is the other half. Here is the translation we use internally, from what the members look like to what a reader should actually do:
| What the members look like | What it means | What to do |
|---|---|---|
| Tightly bundled through day five | The flow is strongly forced and highly predictable | Commit — book the shoot, buy the ticket, dress for the number |
| Tight through day three, fanning after | Confidence has a clear horizon | Make near-term plans firmly and hold the weekend loosely |
| Two distinct clusters | A bimodal forecast with a misleading mean | Plan for both branches and watch for the split to collapse |
| Broad, unstructured fan at day two | A genuinely low-predictability regime | Treat any single forecast as one draw; carry the layer, pack the shell |
| One member far outside a tight pack | A low-probability branch worth monitoring | Track it across runs, but plan around the pack |
Treat a single outlier member as information. One member showing a severe outcome at day six is a low-probability branch worth tracking across runs, though never the scenario you plan a week around.
Spread Grows With Lead Time, And The Rate Is The Signal
Spread widening as a forecast extends is normal, expected, and by itself tells you nothing. Every ensemble looks tight at hour twelve and looks like an exploded haystack at day ten.
What carries information is the spread relative to what is typical at that lead time. A day-two plot as loose as an average day-six plot is a loud statement about the atmosphere's current state.
Wide short-range spread usually means the flow itself is weakly forced — a decaying block, a cutoff low drifting without steering, a stalled boundary that could settle ten miles either way. Under those regimes, skill collapses well inside the range where we normally expect a forecast to hold up.
Narrow long-range spread deserves the opposite reaction. Fifty-one members agreeing on a trough position at day eight means a deeply amplified, strongly forced pattern, and that is the rare case where a week-out forecast is worth acting on.
Compare spread to what is normal for that lead time. Wide spread at day seven is routine, while wide spread at 24 hours means the flow is weakly forced and short-range skill is unusually poor.
Ensemble Plumes: Spread At A Single Point
The spaghetti plot answers spatial questions, and the plume answers point questions. A plume puts time on the horizontal axis, one variable on the vertical, and draws a single trace per member for one specific latitude and longitude.
Temperature plumes are the easiest place to start. A three-degree spread at day three means the number in your app is trustworthy, and a fifteen-degree spread at the same lead time means the app is showing you the middle of a range it has chosen not to mention.
An ensemble plume draws one member trace per line for a single location over time. A three-degree temperature spread at day three is trustworthy; a fifteen-degree spread means your app is hiding a range.
Precipitation plumes behave differently, because rainfall is bounded at zero and unbounded above. Most members cluster near nothing and a handful run very high, which pulls the mean above the median and makes the mean a poor summary of what to expect.
Read the median and the percentile bands on a precipitation plume instead. Our own packing rules for a trip lean on the tenth and ninetieth percentile members rather than the mean, because luggage decisions are decisions about the tails.
Where Ensembles Are Weakest
Reading spread honestly means knowing what spread cannot see. Three failure modes matter enough to change how you interpret a beautifully tight bundle:
- Shared model error. Every member of one ensemble runs the same physics, so a systematic bias is completely invisible in the spread. If a model routinely burns off a marine layer two hours early, all fifty-one members burn it off two hours early and the plot looks confident.
- Underdispersion. Operational ensembles tend to run slightly overconfident, with verified outcomes falling outside the member envelope more often than the envelope implies. The effect is worst for surface variables in complex terrain, where model resolution cannot represent the valley you actually live in.
- Resolution. Ensembles run coarser than their deterministic siblings for the obvious reason that fifty-one runs cost fifty-one times as much. A global ensemble will not resolve a sea breeze front, a lake-effect band, or an individual thunderstorm cell, so its agreement about those features means very little.
All of this adds up to a working caution. Spread is the best public confidence signal available, and it is a correlation established across thousands of cases rather than a guarantee in yours.
The working rule. Spread tells you how much of the uncertainty the model knows about, and model bias is the uncertainty it does not know about.
When the members agree and local experience says the model always busts this particular setup, trust the local experience. Confidence and accuracy are related, and they are not the same measurement.
How We Use Spread In The Daily Brief
Most weather products hide spread because spread is inconvenient to design around. A single number renders cleanly in a widget, and a distribution does not.
We think that trade is backwards. Apple Weather's next-hour precipitation graph — inherited from Dark Sky before Apple retired that app at the start of 2023 — is genuinely good probabilistic design, and it stops at the one-hour horizon where uncertainty is smallest and least interesting.
So when the members disagree, we say so in the brief, in the sentence rather than in a chart. A day where the ensemble splits gets written as a day where the ensemble splits, with both branches named and the decision rule attached.
That is the whole editorial position, and you can see the mechanics of it in how we write a daily brief. A reader who is told confidence is low today makes a better decision than a reader handed a confident number that happens to be wrong.
Common Questions
How many members do the major global ensembles run?
NOAA's GEFS runs 31 members, the ECMWF ensemble runs 51, and the Canadian global ensemble runs 21, counts that have risen steadily as computing costs have fallen. More members sample rare outcomes better, though resolution matters just as much: every ensemble runs coarser than its deterministic counterpart because cost scales with member count.
What is the 500 hPa height line that spaghetti plots usually show?
The 500 hPa geopotential height is the altitude at which air pressure falls to 500 millibars, roughly the middle of the atmosphere by mass, and plots typically contour the 5,460-meter and 5,640-meter lines. Those lines trace the ridges and troughs that steer surface systems, so agreement at 500 hPa is agreement about the large-scale pattern driving your weather.
Does tight ensemble spread guarantee the forecast is right?
No, because every member shares the same model physics, which means a systematic bias produces confident agreement about the wrong answer. Operational ensembles are also mildly underdispersive, so reality falls outside the member envelope somewhat more often than the spread implies — tight spread raises the odds considerably without promising anything.
What is the difference between ensemble spread and run-to-run consistency?
Spread measures disagreement inside a single run at one moment, while run-to-run consistency measures whether successive cycles keep landing in the same place. They are independent signals, and the strongest forecasts have both: a tight bundle that has stayed tight across three or four consecutive runs, rather than one that jumped two hundred miles overnight.
Where can I look at ensemble plots without a subscription?
NOAA publishes GEFS ensemble products through its National Centers for Environmental Prediction model pages, and several free public model viewers render spaghetti plots, plumes, and probability fields for both GEFS and the ECMWF ensemble. Start with 500 hPa height spaghetti for pattern confidence, then move to a point plume for the specific variable you care about.
How should a photographer use ensemble spread when planning a shoot?
Check the cloud cover plume for your exact coordinates rather than the city forecast, and pay closest attention to timing spread. Members that agree on clearing but disagree by six hours about when it happens will ruin a golden hour plan, so keep the date flexible whenever day-three timing spread is wide.