Thursday, 5 September 2013

Spatial Correlation Structures for Different Variables

What and why?                                                                                      
While working on how to consistently homogenise the HadISD data into various different monthly humidity variables (relative humidity RH, vapour pressure e, specific humidity q, wetbulb temperature Tw, dewpoint temperature Td and also drybulb temperature T) I noticed that different stations were kicked out by the PHA homogenisation code.

Main Conclusions: 
It looks like these stations do not contain sufficient correlating neighbours to create a neighbour network. As such they cannot be homogenised by the PHA and so the code removes them from any further processing. There must be at least 7 neighbours that correlate with an r value >=0.1.

It is interesting that different stations are kicked out for different variables:

T:
854690 -27.17S -109.43W 69m ISLA DE PASCUA, Chile (same as Tw, RH)
910660  28.20N -177.38W 4m MIDWAY ISLAND NAS, USA
919250 -9.80S -139.03W 53m ATUONA, French Polynesia
 

q:
085010  39.45N -31.13W 29m FLORES (ACORES), Portugal  (same as Td, Tw)
689060 -40.35S -9.88W 54m GOUGH ISLAND, Tristan de Cunha (same as Td, Tw)
895710 -68.58S 77.95E 13m DAVIS, Antarctica (same as Td)

Td:
085010  39.45N -31.13W 29m FLORES (ACORES), Portugal (same as q, Tw)
689060 -40.35S -9.88W 54m GOUGH ISLAND, Tristan de Cunha (same as q, Tw)
895710 -68.58S 77.95E 13m DAVIS, Alaska

Tw:
085010  39.45N -31.13W 29m FLORES (ACORES), Portugal (same as q, Td)
689060 -40.35S -9.88W 54m GOUGH ISLAND, Tristan de Cunha (same as q,Td)
854690 -27.17S -109.43W 69m ISLA DE PASCUA, Chile (same as RH)

RH:
681040 -22.88S 14.43E 0m WALVIS BAY, Namibia
854880 -29.92S -71.20W 146m LA SERENA, Chile
859300 -52.40S -75.10W 52m FARO EVANGELISTAS, Chile
890220 -75.50S -26.65W 30m HALLEY, Antarctica
919430 -14.48S -145.03W 3m TAKAROA, French Polynesia
689940 -46.88S 37.87E 21m MARION ISLAND, South Africa
847820 -18.05S -70.27W 458m TACNA, Peru
854690 -27.17S -109.43W 69m ISLA DE PASCUA, Chile (same as T, Tw)
854700 -27.30S -70.42W 290m COPIAPO, Chile
916100    1.35N 172.92E 4m TARAWA, Kiribati
916430   -8.53N 179.22E 2m FUNAFUTI NF, Tuvalu

For the most part these are island stations or very isolated locations. An exploration of the different spatial correlation structures of the variables shows that T and Tw have the highest spatial correlation, followed by Td and then q. RH has by far the poorest spatial correlation between neighbouring stations. See Figs 1:5 below. The 40 highest correlating stations (neighbour network) differs for each variable too. This helps to explain both the different station removals and the fact that many more stations are kicked out for RH.

Remaining Issues: 
This means that when thinking about applying adjustments to station data in the process of homogenisation it may be wise to use the neighbour network for that variable. The plan is to use break locations from T and Td (or Tw) and then derive adjustments from variable specific neighbour networks. In some cases the different neighbour network may result in no break being present - in which case, very little adjustment should be applied.

These can be compared with a straightforward direct homogenisation of each variable which will help to quantify the uncertainty stemming from these methodological choices. 

NB: I also found that the PHA correlation output for each neighbour network contained a few instances where a station appeared in its own neighbour network. In one case the station appeared twice. This is not unique to poorly correlating neighbour networks:

For all variables stations: 895320, 918400, 919250 (not T as this station was removed - see above) and 919430 are duplicated within their own networks.

For RH and T station 919580 is duplicated within its own network.

I cannot find a reason for this. It will slightly bias the PHA but with only ~0.001% of stations affected I am not concerned that this has any affect on large scale averages. It may influence the gridbox time series if there are only a few (or only 1) stations within that gridbox. I do not believe this to be a significant problem (see Figs 6:10 below).

Working:
To look at the correlation structure of the neighbour networks I have made a correlation scatter plot showing the candidate station on the x axis and up to 40 of its highest correlating neighbours (must be greater than r=0.1) on the y axis. Each point is coloured by the r value. Stations are identified using the WMO identifier that goes from 000000 to 999999. As a rough guide, European stations go from 000000-199999. Russia and other ex-USSR countries are generally 200000-399999. Most of the Asian countries and Middle-East are 400000-499999. China takes up 500000-599999. African countries take up the 600000-699999. USA has 700000-709999 and 720000-749999 with Canada at 710000-719000. Central America uses 760000-799999. South America uses 800000-879000. Antartica has 880000-899999. 900000-930000 is mostly islands in the Pacific. Australia has 940000-949999. 960000-999999 is for South East Asia.

Figures 1 to 5 show the neighbour network correlation structures for all five variables using the HadISDH.landq.2012p station listings and data download version. The pattern of neighbour networks is very similar but not identical for each variable. The patterns are also not symmetrical because the 40 highest correlating neighbours for one station may not include a lower correlation neighbour if it already has 40 higher correlations. That lower correlation neighbour may be forced to include that station because it does not have sufficient higher correlating neighbours. The density of points in such a small space can also be misleading.

It is clear that T, Tw, Td and q have similar levels of spatial correlation although T and Tw are little higher and q is lower than Td for some networks. RH stands out as having significantly poorer correlations which may explain why so many stations are kicked out for having insufficient correlating neighbours.                  

Figure 1
Specific Humidity neighbour network correlation structure. Black lines indicate stations removed due to too few correlating neighbours (< 7 at > r=0.1).

Figure 2
Relative Humidity neighbour network correlation structure. Black lines indicate stations removed due to too few correlating neighbours (< 7 at > r=0.1).
Figure 3
Dewpoint temperature neighbour network correlation structure. Black lines indicate stations removed due to too few correlating neighbours (< 7 at > r=0.1).
Figure 4
Wetbulb temperature neighbour network correlation structure. Black lines indicate stations removed due to too few correlating neighbours (< 7 at > r=0.1).
Figure 5
Temperature neighbour network correlation structure. Black lines indicate stations removed due to too few correlating neighbours (< 7 at > r=0.1).

Stations that are duplicated within their own network are shown below for T, q and RH to see whether this duplication causes any noticeable problems within the homogenised time series. In the majority of cases there are many neighbours within the network so it is not of great concern. Also, if the adjustments look plausible, as they do in most cases then there is also little cause for concern. The RH adjustments applied for station 918400 are of some concern but this is unlikely due to the duplication of the station within its own network. The PHA adjusts to the most recent homogeneous subperiod which explains the large upward shift to the data above the neighbour network. This has possibly made the station worse, especially for users of absolute values rather than anomalies. With an automated system and zero metadata it is impossible to avoid these issues completely of potentially inaccurate adjustments. However, automation for 3000+ stations is absolutely necessary. One saving grace is that the adjustment uncertainty in this case is also very large so anyone downloading the individual station will see that. On the large scale this individual potential error will have very little impact. On the gridbox scale there will be some impact. In this particular case there are two stations within the gridbox in question so impacts are slightly moderated - average RH is between 75-85 %rh for the gridbox whereas station 918400 averages between 80-85 %rh. In summary, these issues are not thought to be a significant problem for monitoring of recent changes in surface humidity especially for the regional to global scales.               
                                                                      
 Fig 6a q 895320 no adjustments applied
Fig 6b Station 895320 RH
Fig 6c Station 895320 T
Fig 7a Station 918400 q
Fig 7b Station 918400 RH
Fig 7c Station 918400 T
Fig 8a Station 919480 q

Fig 8b Station 919480 RH

Fig 8c Station 919480 T
Fig 9a Station 919580 q
Fig 9b Station 919580 RH
Fig 9c Station 919580 T
Fig 10a Station 919250 q
(no duplication)
Fig 10b Station 919250 RH

Fig 10c T 919250 removed pre-homogenisation






















1 comment:

  1. Interesting. The structure of correlation is somewhat as I would expect. Td was used, of course, in classical synoptic meteorology as an airmass tracer so I'd expect a very high correlation. Because of the non-linearity of C-C I would expect RH to be the lowest correlated as for a linear change in T you get a non-linear change in RH if Td is 'constant'. So, effectively we are doubling down on the variability? Anyway, I would concur that it would make most sense to homogenize T and Td and then build timeseries of inferred variables. The question remains whether Td in particular can be homogenized with seasonally invariant adjustments.

    ReplyDelete

Note: only a member of this blog may post a comment.