Tuesday, 26 November 2013

Applying PHA Indirectly (IDPHA) to RH using PHA Changepoint Locations From T and Td

1) The Plan

Use the locations of changepoints found by PHA for monthly T and monthly Td as changepoints to apply to monthly RH.

Use all stations passing through T and Td PHA successfully (stations fail when they do not have any positively correlating neighbours for that variable - PHA documentation suggests at least 7 neighbours for which r > 0.1 are required but code apparently accepts fewer neighbours).

Use the specific neighbour network defined by station to station correlations of the first difference series of monthly RH anomalies. Neighbour networks differ between variables - see post. Some stations will be removed because there are insufficient correlating neighbours. 

Calculate the difference in the medians of the Homogeneous Sub-Periods (HSPs) in RH anomalies before and after the changepoint for each candidate minus neighbour pair.

Use the median of these pairwise differences as the adjustment to apply to each HSP in RH anomalies. When all adjustments have been applied the mean of the RH anomaly time series will no longer be zero. Remove this offset and then also remove the same amount from the new absolute values once the old climatology has been added back in.

Use the 5th and 95th percentile of all pairwise differences for an HSP to estimate the ~1.65 sigma (90% uncertainty) by using the mean of the 50th - 5th and 95th - 50th percentiles.

Calculate the total adjustment and uncertainty for each HSP whereby the earliest HSP incurs Nchangepoints adjustments, the next HSP incurs Nchangepoints-1 adjustments etc. This is all relative to the most recent HSP which is assumed to be the 'true' climate and so treated as the reference with which to adjust all other HSPs to. To sum the uncertainties the SQRT(sum(uncs^2)) is used.

At the beginning and end of records T and Td may have missing data due to PHA removing data when a robust adjustment cannot be found. This would result in no adjustment being applied to the other variables and no data removed - so potentially a large inhomogeneity remains in the data with no changpoint located e.g., station 042310. So, in cases where the start point of T or Td is not 1, and is greater than 24 months later (cannot be a changepoint detected within first or last two years) then an extra changepoint is added in PHA_indirect at late_startpoint-1. Similarly, if the end point is greater than 24 months earlier than the true end of series (480 months in a 1973-2012 record) then this is treated as a changepoint and an extra point added to contain the true end of series (e.g., 480). If the end point is not 480 and within 24 months then there must be real missing data so  this changepoint is then removed and given an end of series value (480) for computational reasons. 

Adjustments are applied when there is at least one difference series to calculate the adjustment - the changepoint has already been assigned to the candidate station by PHA_direct so we do not have to worry that with only one neighbour we may incorrectly attribute the changepoint to the wrong station.

For the first and last changepoints, where we have applied an extra one because we suspect PHA_direct data removal issues, it is best to remove the changepoint if in fact the data in the candidate were missing all along - otherwise we would count this as a changepoint with 0 adjustment which will mess up statistics calculated later.

2) Things That May Be a Problem 

Over or under application of adjustments. However, any changepoints in T or Td will both be transported through to RH because both are used in the calculation of RH. This would be the case even if RH is measured directly with an RH sensor because only Td is supplied in the ISD data which must then have been derived from RH and T.

Using a different neighbour network to derive adjustments than initially used to define the changepoint location in monthly T and Td may cause errors. For T and Td PHA only the neighbours showing a difference series changepoint are used to estimate the correct adjustment. For PHA_indirect on RH all sufficiently highly correlating neighbours (up to 40) are used. This will likely result in a larger spread of possible adjustments but we hope that where no changepoint is actually present the estimated adjustment will be very very small. We cannot really test this because we do not know definitive dates and locations of the real changepoints.

The PHA_direct tests adjustments for their significance. Where an adjustment is not significant, the original data are removed because there is a detected inhomogeneity but not enough confidence in what to do about it. These potential changepoints are not listed - so while the data remains in the station records for PHA_indirect, no changepoint is allocated and no adjustment is applied. This is not perfect but these are generally cases where the change is ambiguous. In cases where there is a very large change, the adjustment is usually easily quantified. As such, the negative impact on PHA_indirect should be less than the added value of PHA_indirect over PHA_direct (e.g., small inhomogeneities are accounted for in PHA_indirect because they have a stronger signal in the base (T and Td) variables but are too small to be detected by PHA_direct on RH, physical consistency across variables). The PHA_indirect always applies an adjustment where a changepoint is located - regardless of whether that change can be significantly estimated across all relevant neighbours. This may result in larger uncertainty in the adjustments made to the PHA_indirect but this should (in theory) be well quantified in the adjustment uncertainty estimates which are a spread of all possible adjustments across all individual neighbours.

Note that some stations not passing through PHA_direct for Td or T (or PHA_direct on RH) with sufficient months to calculate a climatology may pass through PHA_indirect because no data are removed over periods where a significantly robust adjustment cannot be found. In PHA_direct all adjustments are tested for significance and data can be removed if no adjustment is settled. 

3) Step by Step Discussion

For RH there are 3688 starting stations that pass through both T and Td PHA out of 3694 potential stations. The following stations are removed:

     WMO/WBAN ID, latitude, longitude, elevation, country, name  
Td 08501099999  39.45    -31.13           29 PO FLORES (ACORES)             
Td 68906099999 -40.35     -9.88           54 ZA GOUGH ISLAND                
T   85469099999 -27.17   -109.43           69 CH ISLA DE PASCUA              
Td 89571099999 -68.58     77.95           13 AY DAVIS                       

T   91066022701  28.20   -177.38            4 US MIDWAY ISLAND NAS           
T   91925099999  -9.80   -139.03           53 PF ATUONA                      


After running the indirect PHA there are now 3675 stations. 13 more stations are removed because they do not have sufficient correlating neighours (<7 where r > 0.1):
WMO/WBAN ID, latitude, longitude, elevation, country, name, No. of Neighbours  
61967070701 -7.3000     72.4000    2.7  BT DIEGO GARCIA NAF                 5 
68104099999 -22.8830   14.4330    0.0  NM WALVIS BAY (PELICAN            0
68994099999 -46.8830   37.8670   21.0 ZA MARION ISLAND                       0

78767099999 10.0000  -83.0500    3.0   CS PUERTO LIMON                       6  
84782099999 -18.0500  -70.2670  458.0 PR TACNA                                     0
85470099999 -27.3000  -70.4170  290.0 CH COPIAPO                                 0
85488099999 -29.9170  -71.2000  146.0 CH LA SERENA                             0
85930099999 -52.4000  -75.1000   52.0  CH FARO EVANGELISTAS             0
89022099999 -75.5000  -26.6500   30.0  AY HALLEY                                   0
91610099999  1.3500  172.9170    4.0   KB TARAWA                                  0
91643080705 -8.5330  179.2170    2.1   TV FUNAFUTI NF                           0
91943099999 -14.4830 -145.0330    3.0 PF TAKAROA                                 0
93945099999 -52.5500  169.1670   19.0 NZ CAMPBELL ISLAND AWS         6


This result differs to PHA_direct which results in only 11 stations failing the neighbour criteria. Station 619670, 786670 and 934450 are new 'bad' stations in PHA_indirect relative to PHA_direct. This suggests that although documentation for PHA_direct states where possible stations should have at least 7 neighbours correlating with r > 0.1, they can pass through with fewer. In PHA_indirect I have set the minimum to 7. Also, station 854690 has already been removed as it did not pass through PHA_direct on T. This accounts for the 13 stations removed here in PHA_indirect as opposed to 11 in the PHA_direct.

The ascii output of adjusted absolute RH is then converted to netCDF files containing anomalies and absolute values. Anomaly values can only be calculated when there are enough (at least 15) months present during the climatology period (1976-2005). During PHA_direct this can remove more stations because some periods of data are removed when a significant adjustment cannot be resolved upon. The PHA_indirect is less sophisticated and always applies an adjustment. However, some stations are still removed during calculation of anomalies because there are insufficient years represented across the climatology period for any one month. This can occur, even though the original code to calculate monthly averages has not picked this up. This is because the original code builds monthly absolute and anomaly values from hourly absolute and anomaly values. There can be a number of month_hour_means present over the years to create a climatology for that month_hour but insufficient month_hours to create a mean_month value for any one particular month. This is a weakness in the code and so should be corrected before the next build of the system (2013).

*** Correcting this error would result in slightly fewer stations being passed through to PHA (or IDPHA) in the first place. This may have the effect of changing a few of the adjustments applied as the neighbour networks can differ slightly. This may then also feed through to the gridbox values, as would the removal of a few stations. For the large scale means this is highly unlikely to have any effect. For this stage of comparison I will leave the code as is as the data have already been processed and reprocessing is time expensive at this stage. I plan to bugfix prior to the 2013 update for all HadISDH variables. ***

In total 30 stations are removed because there are insufficient data to calculate anomalies - these should have been removed before applying PHA_direct/PHA_indirect (see above). Some are listed as zero months - this is when the PHA_indirect code has found adjustments and so had to process the anomalies. This anomaly process failed due to too few months being present leaving the missing data indicator for the whole month. A further 26 stations have instances of RH > 100 %rh which is unphysical at the monthly scale. This leaves 3619 'good' stations.

30 Stations with insufficient monthly means (<15) to calculate a climatology (1976 to 2005):
080480,082320,229390,270080,303790,306370,307120,307290,400610,415940,433330,
612700,614920,620560,621610,637910,648600,678530,804070,804190,812000,860110,
860680,862600,870160,875930,876400,889630,890660,972600 


26 Stations with at least on month being > 100 %rh: 
010620,012280,014060,064760,064900,081410,101420,111050,111460,151080,160840,
200690,200870,202920,230320,320760,470610,547760,548260,563850,584370,723910,
724837,727855,745160,890560

These are not the same as stations with > 100%rh caused during PHA_direct except for the 10 stations highlighted in blue:
061200,064760,110220,110280,111300,151080,200870,202920,219820,230320,319600,
470610,547760,577760,584370,586660,700260,700860,703080,718936,724837,725975,
745160,857990,870970,984400


The stations are then averaged over 5by5 degree boxes to provide monthly mean gridbox anomalies. 

3a) Average number of changepoints applied per station and distribution over time for PHA_indirect verses PHA_direct:

PHA_direct results in 2.6 adjustments per station on average compared to 3.35 per station in PHA_indirect. This includes all 'good' stations - so including those that have no detected inhomogeneities.

There are far more adjustments applied in PHA_indirect in years 1975/1976 and 2009/2010 (Figure 1) than for PHA_direct (Figure 2) which tails off towards either end of the record. Having looked at a number of stations that have changepoints in these peaks a number are real changepoints assigned from T and Td. Another lot are from changepoints assigned at the apparent start or end of the homogenised T or Td records where data have been removed by PHA_direct. In these cases, PHA_indirect does have data to work with and so assigns and adjustment where as PHA_direct cannot robustly assign an adjustment and so removes the data and does not count the potential changepoint. 

The application of changepoints at end points for PHA_indirect that have not been counted in PHA_direct because an adjustment cannot be robustly assessed is at least one reason why there are more changepoints in PHA_indirect that PHA_direct (even compared to a straight comparison of PHA_direct changpoints summed from T and Td which have 1 and 2 changepoints per station found on average, respectively). The summing of changepoints from T and Td PHA_direct will also mean that RH PHA_indirect has more changepoints than RH PHA_direct because it is likely that changepoints will be detected in T and Td where they may not be detectable for RH.

Figure 1
PHA_indirect (RH) adjustment size distribution a) and time distribution b). For a) the black steps are the actual adjustment distribution, the grey line is the best fit Gaussian curve, the red dashed line is a merge of the fat tails in the actual histogram and the missing middle of the Guassian curve and the blue dotted line is the difference between the merged distribution and the actual distribution. The mean and standard deviation of the difference is shown - these are used to estimate uncertainty remaining in the data due to unassigned changepoints.




Figure 2
PHA_direct (RH) adjustment size distribution a) and time distribution b). For a) the black steps are the actual adjustment distribution, the grey line is the best fit Gaussian curve, the red dashed line is a merge of the fat tails in the actual histogram and the missing middle of the Guassian curve and the blue dotted line is the difference between the merged distribution and the actual distribution. The mean and standard deviation of the difference is shown - these are used to estimate uncertainty remaining in the data due to unassigned changepoints.

 

3b) Average size and distribution of changepoints applied compared to direct PHA

PHA_direct results in an average absolute adjustment size of 3.94 %rh on average with 1 standard deviation = 2.38 %rh. For PHA_indirect the average absolute adjustment size is 2.84 %rh with one standard deviation = 2.21 %rh. 

PHA_direct results in an average adjustment size of 0.1 %rh on average with 1 standard deviation = 4.6 %rh. For PHA_indirect the average adjustment size is -0.01 %rh with one standard deviation = 3.59 %rh. 

For both, the mean of all adjustments is around zero. The absolute adjustments are slightly smaller for PHA_indirect compared to PHA_direct. There are many many more very small adjustments applied during PHA_indirect compared to PHA_direct. The signal of changepoints in Td can be much larger than the same changepoint in q and RH - especially in cold/low RH/low q climates. In this way PHA_indirect could be argued to be a better choice because otherwise these small adjustments would be missed.
 

3c) Comparison of PHA 12 largest adjustments (station ID and total adjustment in %rh):

PHA_indirect                   PHA_direct
36982099999  25.76      04231099999  26.3100   
04231099999 -25.75     
36982099999 -25.7900
76577099999 -19.14
      44284099999 -22.9400
72654899999 -18.64     
44277099999 -22.5800
60714099999  17.00     
72654899999  22.2500
44287099999  16.18
      44287099999 -20.6400
59948099999 -16.04
      44213099999 -18.8500
78367011706 -15.84
      36982099999  17.9600
17290099999 -15.69
      38545099999 -17.9300
04202099999 -15.56
      04260099999 -17.2500
44284099999  14.86
      60714099999 -17.0900
44347099999  14.79
      44218099999  16.9900

These are comparable in size with PHA_indirect leading to smaller adjustment magnitudes over all. The differences in stations experiencing the largest adjustment could be due to three reasons. Firstly, PHA_direct removes data over periods where an adjustment cannot be robustly estimated so PHA_indirect, while not including those particular changepoints, will have more data present to compare HSPs and so adjustments will differ slightly. Secondly, PHA_indirect may have changepoints in a slightly different place compared to PHA_direct for the same inhomogeneous feature which would result in different sizes in the adjustments. Thirdly, the neighbour network used to estimate the adjustment size will be different as PHA_direct uses only those neighbours where a changepoint is apparent compared to PHA_indirect which uses all neighbours within the specified network.

All stations in the PHA_indirect list are shown below (Figs 3 to 14) compared with RH from PHA_direct where the station passes the sufficient data to calculate a climatology test (after removal by PHA_direct - no station 443470 for PHA_direct). Raw trends differ slightly because of the different software used to calculate the linear trend using the median of pairwise trends. This will be reconciled at a later date.

Station 369820 - Figure 3a, b:
PHA_indirect returns a much smaller negative trend that is no longer significantly different from a zero trend, compared to PHA_direct. Quite a lot of data are removed in PHA_direct which leave the assessment of any trend ambiguous. PHA_indirect appears to present a much more homogeneous record from 1985 onwards. Whereas PHA_direct may represent the 1980-1985 period better it appears to introduced a changepoint between 2003-2005. Ultimately it is difficult to know which is truly better as we do not know the true record. PHA_indirect does not look worse, and if anything, looks better.

Station 042310 - Figure 4a, b: 
PHA_indirect returns a much smaller negative trend that is no longer significantly different from a zero trend, compared to PHA_direct. Some data are removed in PHA_direct which leave the assessment of any trend a little ambiguous. PHA_indirect appears to present a much more homogeneous record from 1973-1986. Ultimately it is difficult to know which is truly better as we do not know the true record. PHA_indirect does not look worse, and if anything, looks better. Both versions look outside the neighbour range due to adjustment relative to the most recent period.

Station 765770 - Figure 5a, b: 
PHA_indirect returns a much smaller negative trend that is no longer significantly different from a zero trend, compared to PHA_direct. Some data are removed in PHA_direct which leave the assessment of any trend a little ambiguous. PHA_indirect appears to present a much more homogeneous record in general albeit with large variability between 1995-2005. 
Ultimately it is difficult to know which is truly better as we do not know the true record. PHA_indirect does not look worse, and if anything, looks better.

Station 726548 - Figure 6a,b:
PHA_indirect returns a positive trend that is significantly different from a zero trend, compared to PHA_direct which gives a small negative trend that is not significantly different from zero. Quite a lot of data are removed in PHA_direct which leave the assessment of any trend a little ambiguous. The increase shown in PHA_indirect is drive mostly by the rise across the 1994-1997 period and may well be spurious. Ultimately it is difficult to know which is truly better as we do not know the true record. 

Station 607140 - Figure 7a, b: 
PHA_indirect returns a smaller negative trend that is no longer significantly different from a zero trend, compared to PHA_direct. Both methods cope well with the large inhomogeneity across 2000-2010 and look plausible. 

Station 442870 - Figure 8a, b: 
PHA_indirect returns a smaller negative trend that is not significantly different from a zero trend, compared to PHA_direct which returns a very small positive trend that is not significantly different from zero. Some data are removed in PHA_direct which leave the assessment of any trend a little ambiguous. Both improve the very inhomogeneous raw station record. Ultimately it is difficult to know which is truly better as we do not know the true record. 

Station 599480 - Figure 9a, b: 
Here both methods perform very similarly resulting in small negative trends (~0.4 %rh per decade) that are significant - with slightly larger spread for PHA_indirect. Both appear to improve this inhomogeneous record. Ultimately it is difficult to know which is truly better as we do not know the true record. Both versions look outside the neighbour range due to adjustment relative to the most recent period.

Station 783670 - Figure 10a, b: 
PHA_indirect returns a small positive trend that is not significantly different from a zero trend, compared to PHA_direct which returns a small negative trend that is significantly different. Some data are removed in PHA_direct which leave the assessment of any trend a little ambiguous. The decrease across the 1985-1990 period in PHA_direct looks a little spurious but then so does the variability in PHA_indirect pre-1995. Ultimately it is difficult to know which is truly better as we do not know the true record. PHA_indirect does not look worse, and if anything, looks better. 

Station 172900 - Figure 11a, b: 
PHA_indirect returns a larger negative trend that is now significantly different from a zero trend, compared to PHA_direct. Quite a lot of data are removed in PHA_direct which leave the assessment of any trend a little ambiguous. The period from 1985 onwards is similar for both methods. PHA_indirect retains far more data which may cause the spurious decline between 1973 and 1985. Ultimately it is difficult to know which is truly better as we do not know the true record.

Station 042020 - Figure 12a, b: 
PHA_indirect returns a notable positive trend that is significantly different from a zero trend, compared to PHA_direct which shows no trend. Quite a lot of data are removed in PHA_direct which leave the assessment of any trend a little ambiguous. The variability in the earlier record from PHA_indirect (and PHA_direct although some data are removed) is much greater than the variability in the later record. The increasing trend here may well be spurious but the record from PHA_direct does not look much better. Ultimately it is difficult to know which is truly better as we do not know the true record.

Station 442840 - Figure 13a, b: 
PHA_indirect returns a near-zero positive trend that is not significantly different from a zero trend, compared to PHA_direct which gives a notable and significant negative trend. Quite a lot of data are removed in PHA_direct which leave the assessment of any trend a little ambiguous. Both methods perform well at removing the large inhomogeneity - the negative trend resulting in PHA_direct may be spurious. Ultimately it is difficult to know which is truly better as we do not know the true record. PHA_indirect does not look worse, and if anything, looks better.

Station 443470 - Figure 13a:
Insufficient data passed through PHA_direct for this station to become a homogenised station record in HadISDH. For PHA_indirect the homogenised record still looks ambiguous with an apparent large inhomogeneity remaining in the first year and large variability throughout. However, the large inhomogeneity in the middle of the record is mostly accounted for by the adjustments. The resulting trend, although negative, is not significantly different from a zero trend.

Summary:
In my view PHA_indirect performs ok. While not significantly outperforming PHA_direct in these cases it does retain more data and consistency with the other humidity and temperature variables.

Figure 3a
PHA_indirect homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 369820.
 
Figure 3b
PHA_direct homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 369820.
Figure 4a
PHA_indirect homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 042310.

Figure 4b
PHA_direct homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 042310.

Figure 5a
PHA_indirect homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 765770.

Figure 5b
PHA_direct homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 765770.

Figure 6a
PHA_indirect homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 726548.
Figure 6b
PHA_direct homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 726548.

Figure 7a
PHA_indirect homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 607140.
Figure 7b
PHA_direct homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 607140.


Figure 8a
PHA_indirect homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 442870.
Figure 8b
PHA_direct homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 442870.
Figure 9a
PHA_indirect homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 599480.
Figure 9b
PHA_direct homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 599480.
Figure 10a
PHA_indirect homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 783670.
Figure 10b
PHA_direct homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 783670.
Figure 11a
PHA_indirect homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 172900.
Figure 11b
PHA_direct homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 172900.
Figure 12a
PHA_indirect homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 042020.
Figure 12b
PHA_direct homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 042020.
Figure 13a
PHA_indirect homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 442840.
Figure 13b
PHA_direct homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 442840.
Figure 14a
PHA_indirect homogenised (blue) verses raw (red) annual average RH against raw neighbours (black) with decadal trends (median pairwise) and 5th-95th percentile confidence intervals for station 443470.

Figure 14b - no PHA_direct figures available due to failing to pass test for having sufficient numbers of months after PHA for calculating a climatology.

 

3d) Spatial comparison of adjustment sign and magnitude between PHA_direct and PHA_indirect for RH:

The latitudinal distribution of adjustments and their magnitude is very similar between PHA_direct (Figure 15) and PHA_indirect (Figure 16). Although there are many many more small adjustments made in PHA_indirect the symmetry across positive and negative adjustments is common to both and the larger adjustments in the mid-latitudes. Clearly there will be more adjustments where there are more stations i.e., the mid-latitude Northern Hemisphere.

There are some differences in the magnitude and sign of the largest adjustments made at each station (compare parts b) of both figures). For example, PHA_direct suggests larger positive adjustments and the Scandinavian region is mostly blue (positive adjustments) in PHA_direct whereas it is mostly brown (negative adjustments) in PHA_indirect. Some of this could be due to the fact that more smaller adjustments are applied in PHA_indirect such that some large changes are adjusted incrementally as opposed to in one adjustment whereas in PHA_direct, due to non-adjustment over non-robust adjustment estimates and the inability to detect smaller changes, there are fewer larger adjustments (see Section 3b).


Figure 15
a) Adjustment size by latitude and b) maximum adjustment size for each station for RH PHA_direct.

Figure 16
a) Adjustment size by latitude and b) maximum adjustment size for each station for RH PHA_indirect.

 

3e) Comparison of resulting decadal trends from PHA_indirect compared to the raw data:

Figure 17 shows that for the most part PHA_indirect is bringing trends closer to zero compared to the raw data. The majority of gridboxes (75 %) retain the same direction of trend after homogenisation but a considerable proportion (25 %) change the direction of their trend. For regional averages, in all cases (Globe, NH, Tropics, SH) the trends are moderated by homogenisation. While the decline post 2000 is still very apparent, declining trends are not significant for the Globe or Tropics, where they were in the raw data. The Southern Hemisphere shows significant declining trends in both the raw and homogenised data.

Figure 18 shows the decadal trends in RH for PHA_indirect. Drying out/negative trends are widespread and spatially coherent for the mid-latitudes in both hemispheres (over large land masses, not islands) and moistening trends are apparent over the Tropics and high latitudes.
Figure 17
Comparison between raw and PHA_indirect homogenised RH. a) ratio of gridbox decadal trends 1973-2012. b) Scatter plot of decadal trends for each gridbox with percentage of gridboxes in each sector. c) Histogram of decadal trends for all gridboxes. d) Large scale area average time series and decadal trends (median pairwise) with 5th-95th percentile confidence intervals.
Figure 18
Decadal trends in monthly mean RH from PHA_indirect from 1973 to 2012. Trends are fitted by the median of pairwise slopes. Where both the 5th and 95th percentiles of the slopes are on the same side of zero a trend is considered to be significant, this is identified by a black dot.

3f) Comparison of resulting decadal trends between PHA_indirect and PHA_direct for RH:

The difference in decadal trends between PHA_direct and PHA_indirect is actually on a par with the difference between the raw and homogenised data for PHA_indirect. Importantly, the drying out is stronger in PHA_direct and significant for the Globe and the Northern Hemisphere where as it is not significant for PHA_indirect. There is good agreement over the Tropics. Over the Southern Hemisphere, PHA_direct results in a weaker trend that is not significantly different from a zero trend where as PHA_indirect gives a strong significant trend (Figure 19d).

At the individual gridbox scale PHA_indirect results in slightly smaller (closer to zero) trends over all. For 30% of gridboxes the trend changes sign between PHA_direct and PHA_indirect (Figure 19a,b,c).

Figure 20 (compared to Figure 18) shows that widespread drying is common to both methods. However, in PHA_indirect this does appear more latitudinally discrete with a more solid picture of drying over the mid-latitude Southern Hemisphere and moistening over the Tropics and high latitude Northern Hemisphere.


Figure 19
Comparison between PHA_direct and PHA_indirect RH. a) ratio of gridbox decadal trends 1973-2012. b) Scatter plot of decadal trends for each gridbox with percentage of gridboxes in each sector. c) Histogram of decadal trends for all gridboxes. d) Large scale area average time series and decadal trends (median pairwise) with 5th-95th percentile confidence intervals.

Figure 20
Decadal trends in monthly mean RH from PHA_direct from 1973 to 2012. Trends are fitted by the median of pairwise slopes. Where both the 5th and 95th percentiles of the slopes are on the same side of zero a trend is considered to be significant, this is identified by a black dot.

 

4) Conclusions

A decision has to be made whether to use PHA_direct or PHA_indirect, on the assumption that PHA on the whole is a suitable homogenisation method. 

The overall result is reasonably similar in that there is a signal of widespread drying. However, this appears to be more latitudinally discrete in PHA_indirect with a weaker (insignificant) signal over all when considered as a long-term trend over the whole period.

Consistency of adjustments across humidity and temperature variables is very important, especially to studies utilising both components. PHA_indirect provides the strongest chance of ensuring this compared to PHA_direct. It has also been demonstrated that very small changes to q and RH are not easily detected when searched for directly but are easier to detect from the base variables of T and Td where signal-to-noise ratio is better. This also provides support for PHA_indirect over PHA_direct.

One downside of using PHA_indirect is that it loses some of the sophistication of the PHA code in terms of robustly assessing adjustments only with those neighbours where a changepoint is apparent and only making the adjustment when it is considered to be significant. This is not something that can be easily brought in to the PHA_indirect code. Changes are more difficult to detect so it would be hard to define the exact neighbour network within which the changes are apparent. Also, neighbour networks differ slightly between variables (see here). This means that distributions of potential adjustments considering all neighbours are likely to be wider spread AND include more very small changes relative to the distribution for PHA_direct - so potentially harder to quantify as robust or not?

Overall, my decision would be to apply PHA_indirect. There are occasions where PHA_direct may perform better on individual stations, especially where data are removed in instances of great uncertainty. However, given that the underlying truth is unknown it is really very difficult to choose one over the other on eyeball analysis of the data. The consistency of adjustments across variables and ability to detect small changes are the winning arguments for me.


 

No comments:

Post a Comment

Note: only a member of this blog may post a comment.