The argument behind the daily standings, and the standings themselves. The day-by-day record they are drawn from is on the scorecard.
The standings published daily measure overall forecast accuracy across ForecastEx and eleven public forecast tools, averaged over a rolling 7-day window with a matched sample across every system.
This constitutes ongoing evidence that weather prediction markets like ForecastEx may be the most accurate short-range weather forecasts available, and the difference becomes greater as lead time decreases.
The exchange against LAMP, the National Weather Service’s observation-updated aviation guidance, which is the one public product that reissues often enough to be compared at every lead. Each system’s forecast of the daily high against the temperature the station recorded, pooled across every city and every hour of the day, plotted by how long before the end of the day it was made. Hover a point for the two errors and the number of city-days behind it. The other three systems this site tracks issue a few times a day and are compared on the scorecard.
Over the last week of scored days, how far was each tool from the temperature the station recorded? Ranked best first, on a matched sample so every tool is judged on the same station-days. Hover a bar for its bias, its share within two degrees, and the same figures on the daily low.
Both systems are scored the same way, on the same days, against the same recorded high. The gap is small a day and a half out, where neither system knows much the other does not, and widens as the day fills in, because the market can price an afternoon that is already half observed.
Daily weather forecasting is one of the most mature, established, and scientifically principled fields of science and industry. It is not an exaggeration to describe conventional weather forecasting systems as the frontier of applied physics, statistics, and computer science, resting on decades, if not centuries, of scientific inquiry.
As such, conventional weather forecasts should hardly represent a target ripe for a novel forecasting system to improve upon. Despite the formidable challenge, ForecastEx prediction markets seem to be doing just that, not just echoing public forecasts but aggregating dispersed information in a way that improves upon sophisticated established systems.
This is possible because prediction markets are not substitutes for standard forecasts but rather sit downstream of them. They incorporate conventional forecast information as one input and convert it into refined probabilities through direct financial rewards for being accurate and direct financial penalties for being inaccurate. This creates a dual effect of attracting accurate individuals and systems into the market while deterring those who are inaccurate. People or systems that consistently make poor forecasts are heavily motivated to either improve or leave the market.
The argument here is about market prices as forecasts in general. The scorecard carries the measured record, including the days the market did worse than the forecast products, scored on a matched sample so every system is judged on the same days. Nothing here is investment advice.