Every active station currently in the grid (or just the network you've picked above), plotted at its real coordinates. Scroll to zoom, drag to pan, click a point for its latest reading. CoCoRaHS alone is ~51,000 stations — filter down to nws/ndbc/metar to actually see the temperature-reporting stations instead of a screen full of rain gauges.
| Station | Name | Network | Last Obs (UTC) | Temp °F | Wind mph | Gust mph | Precip in | RH % | Pressure inHg |
|---|---|---|---|---|---|---|---|---|---|
| Loading… | |||||||||
A handful of these is normal — some rural or decommissioned stations don't report every field, or don't report at all anymore. A wall of errors, or zero observations stored anywhere above, would be the sign something's actually wrong.
Two different questions, both answered from the same ground-truth grid. Forecast Model Accuracy is the real point of a "Benchmark" — the same gap each app already closes at its one station (a real reading vs. a general model's prediction), measured here across a sample of stations nationwide. Ground-Truth Consensus is a supporting data-quality check on the grid itself — whether nearby stations agree with each other — not a forecast comparison.
Not a nationwide model comparison — this is each app's actual production forecast (every source blended, weighted by track record, bias-corrected, and self-corrected against its own tracked residual), scored week by week against whichever single model was most accurate that same week, at that app's own real ground-truth station.
Open-Meteo's GFS, ECMWF, and ICON models — free at this scale, unlike AccuWeather/Meteoblue, which stay scoped to each app's one primary station — forecasting for a sample of Benchmark stations nationwide, graded against what those stations actually recorded once the date arrives. Temperature, wind, and precipitation are graded independently, each in its own native unit (°F average high/low miss, mph daily-max miss, inches daily-total miss). A model only gets one new graded day per station per metric, so … graded days is when a score is called Tracked here (a lower bar than Ground-Truth Consensus's, since samples arrive far less often). On each card: the big number is that model’s average miss against ground truth in the unit shown — lower is better — averaged over the Tracked stations only, which is what the count beneath it reports.
Pick one model above (not "All models") to see its error blended across the map — one continuous green-to-red scale per station, not a handful of buckets, so you can see at a glance where a model runs sharp versus where it struggles. Dimmer dots are still "Building" (too few graded days to trust yet); full-opacity dots are "Tracked." Uses the same model/metric/lead-time picked for the table below, but always plots up to 2,000 stations nationwide regardless of the table's own page size, so the map never looks sparser than the grid actually is.
| Station | Name | Network | Region | Model | Avg Error | Graded Days | Status | Last Scored |
|---|---|---|---|---|---|---|---|---|
| Loading… | ||||||||
How well each station agrees with its nearby peers — a data-quality check on the ground-truth grid itself, not a forecast comparison. On each card: the big number is the average gap between a station and its nearby neighbours — lower means tighter agreement — and it is averaged over the Tracked stations only, which is exactly what the count beneath it is telling you. Every collector run adds one more comparison sample per station, so a score marked "Building" firms up over time; … samples is when it's called Tracked.
| Station | Name | Network | Region | Avg Deviation | Samples | Status | Last Scored |
|---|---|---|---|---|---|---|---|
| Loading… | |||||||