Browser weather
Do you feel like the weather models aren’t very good at predicting temperature near you? That’s because in most cases their closest data points are the airports, far from you.
Airports report the weather every half hour or so, in a short coded format called METAR: temperature, dew point, wind, pressure, visibility, clouds. A report looks like LFPG 231200Z 24012KT 9999 SCT030 18/11 Q1015. These reports are public, and some services serve them in a way a web page can read directly: the Iowa Environmental Mesonet for the whole world, the US National Weather Service for the US. Everything below is fetched and computed by your browser.
Pick a place:
choose a location
So we have weather station data. As you can probably see, most of the stations are far away from you.
One very naïve thing we can do, if we want to know the temperature where we are, is to average those values.
A simple improvement we can make is to take into account how far the stations are. The stations close to you should count more than the ones far away, so we weight each one by the inverse of its distance: a station twice as far counts half as much.
The far ones barely count: a station 50 km away weighs a tenth of one 5 km away. So there is no point in fetching many of them. We take the 5 nearest within 100 km, which also keeps the data small.
Here are the 5 stations used for your place, and the share of the total each one gets:
We get a first estimation, and it costs us ms to fetch the reports, then to compute.
Now, we only take into account the distance, but other things factor in. Say you live at the top of a mountain and there’s an airport down in the valley, just a few kilometers away. You probably have pretty different temperatures, even though you’re not that far away.
So we correct for a few things:
- Temperature drops with altitude, so each station’s temperature is brought to the altitude of your point, using the temperature gradient measured across the stations themselves, or the textbook 6.5 °C per kilometer when there aren’t enough stations to measure it.
- Dew point doesn’t average well as it is: the amount of water in the air grows exponentially with the dew point, so the average of two dew points isn’t the dew point of the mixed air. We average the water vapor pressure instead, which follows the amount of water, and convert back.
- Pressure is averaged as the sea-level value the airports report, and kept at sea level, like theirs.
- Wind is averaged as direction and strength together, so a north wind and a south wind cancel out instead of averaging to an east wind.
Under each value, in grey, is what a weather model says for the same point right now, from Open-Meteo, which picks the best model available for your area:
This is still a naive approach. Weather models are fed with data from these same stations, amongst many others (radar, satellites, weather balloons), and run physical models of the atmosphere to compute a value for your location. This one needs a handful of reports and a fraction of a millisecond, runs entirely in your browser, and, as the grey lines show, can be pretty close to what they get.
P.S.: if you want to go further, here are three methods that climatologists use to fill the gaps between stations:
- GIDS (“gradient-plus-inverse distance squared”, Nalder and Wein, 1998), which “combines multiple linear regression and distance-weighting”. A slightly more complex version of what we do here. It fits how temperature changes with longitude, latitude and altitude across the nearby stations, then weights them by the inverse square of their distance. Still cheap in data and compute: a few stations and a small regression per point.
- Thin-plate smoothing splines (Hutchinson’s ANUSPLIN, used for the WorldClim climate maps). It fits one smooth surface through all the stations at once, with altitude as an extra dimension, and chooses how smooth by cross-validation. Smoother and better grounded statistically. But it needs every station at once: the exact fit solves one system as large as the number of stations, fine on a computer once per update, heavy in a phone’s browser. It also smooths away local effects.
- PRISM (Parameter-elevation Regressions on Independent Slopes Model, Daly and colleagues, Oregon State University). It fits a temperature-versus-altitude relation for every map cell, weighting stations by distance, altitude, which side of the mountain they’re on, and distance to the coast. The most accurate in mountains, and behind the official US climate maps. But it needs a detailed elevation model and a lot of expert tuning, and it’s built for long-term averages, not the weather right now.
You are probably wondering how good this is on more than just the one location you just tried.
For each station, we mask its own value, keep it as ground truth. We then run our naive method here at the Ground Truth Station location we just masked.
Doing this for all stations (the ones with fresh reports, and enough neighboring stations), we can measure:
The average error was 1.19 °C. What drives that error is mostly the height gap to the nearest neighbour, not the distance: stations whose nearest neighbour is within 100 m of their own elevation average 1.02 °C off, while stations whose nearest neighbour is more than 1,000 m higher or lower average 3.36 °C off.
There are a few ways we could improve our approach here:
- making sure the fit can’t run away, by limiting our correction rate to values that have a physical sense.
- fitting those rates on more stations.
- Also weight the neighbours by height difference as well as distance.
Reuses the kNN, IDW, correction and source-fetching modules from /knn-weather/. Stations: idle-intelligence/metar-stations, built from Iowa Environmental Mesonet rosters. Observations: Iowa Environmental Mesonet, api.weather.gov; elevation and weather model values from Open-Meteo.
Source: github.com/idle-intelligence/weather-web, compiled to WebAssembly for this page.