The authors have performed several new analyses in order to respond to my admittedly lengthy list of requests, and I appreciate their efforts. While I think another round of major revisions is required, I do think this manuscript is on its way to being acceptable for publication. In this review, line numbers refer to the 'track changes' document.
First, I will respond to a few points in the authors' response:
"Spatial overlap of known limitations and observed performance:
The reviewer’s interpretation rests on the assumption that poor model performance in Figure 5 should spatially coincide with the regions of known GLOBGM limitations flagged in Figure 4. We respectfully challenge the validity of this assumption..."
My interpretation was actually somewhat more general: if a paper states that results are less reliable in some regions (and hence more reliable in some others) I expect that the high-reliability regions display better performance. If the results are in fact quite comparable between the regions, why delineate regions at all? Or what limitations have not been considered that are reducing performance in the high-reliability regions?
"Then regarding the suggestion to focus only on regions where the model matches historical observations we respectfully disagree with this approach. Groundwater monitoring networks are heavily concentrated in the Global North (Bäthge et al. 2026) and in addition groundwater models are overwhelmingly concentrated in high-GDP countries with sparse coverage over Africa, South Asia, and South America (Zamrsky et al. 2025). Restricting our model output to regions with good simulation and observation agreement would systematically exclude the Global South. We argue that providing physically-based model estimates, while acknowledging uncertainty, is more informative than providing nothing for these regions."
I appreciate the point, but there is a fine line between reporting highly uncertain results and reporting results that we know are probably wrong, without clearly stating this. The regions the authors point out that have sparse coverage (Africa, South America, South Asia) contain substantial areas where the model disagrees strongly with GRACE observations. If the model is wrong in these locations, releasing these results is probably worse for decision-makers than releasing no results at all. The manuscript is on its way to a better discussion of uncertainty, and I think that I will be comfortable with publication with more clear and upfront statements about the model failures.
GRACE comparison: Thanks for including this. However, more needs to be said about the comparison.
- For what percentage of the global land surface does GLOBGM show the same direction of change as GRACE?
- For what percentage of the land surface do the two estimates agree within some margin of error?
- Is poor meteorological forcing really the only cause of this poor performance? There's a wealth of research on groundwater and recharge in sub saharan Africa, and a couple of papers are cited, but surely some more of it could be useful? What about shifts in precipitation intensity? eg. Taylor et al (2013) https://doi.org/10.1038/nclimate1731, Zhang et al (2016) https://doi.org/10.1002/hyp.10809, Jasechko and Taylor (2015) https://doi.org/10.1088/1748-9326/10/12/124015. Is heavy-rainfall-driven recharge well-represented in GLOBGM? Perhaps there are other important processes that are also missing from the model. I don't think these issues need to be corrected in the model for this publication, but I do think that they need to be pointed out so readers can judge for themselves whether they trust the results for a particular region.
In addition to sub-saharan Africa and Siberia, I think it's worth pointing out other areas where the model does not match GRACE:
- Peninsular India
- the Alps
- The Andes
- The Caspian Sea and south of the Black Sea
- The North American Great Lakes
- North-western North America
- Mexico
- Northern South America
- southeast US
- northeast US and eastern Canada
- Japan
- Iceland
- New Zealand
- Australia is also quite dissimilar if you look at it region-by-region.
- I don't agree that the simulations for northern Africa agree with GRACE. Perhaps for the Sahara desert, but not for the populated coastal regions.
The authors have improved the discussion of the model performance, but the model limitations (for example, for the above regions) should also be made clear in the abstract.
L497: For these updated performance scores, am I correct that the authors only screened out locations where the following criteria are met? ((the observed and simulated trends are <0.1 m) AND (rs<0)) OR ((the observed and simulated heads are <1 m) AND the KGENPskill<0)? Obviously, if you discard only the poor simulations (rs<0 or KGENPskill<0) then the overall skill improves! If the authors want to set some criteria (based on heads/trends) to filter results before evaluating performance, then they also must discard skilful predictions that meet those criteria. This is a basic principle of model evaluation.
What is the justification for using three different thresholds to mask out non-significant or small trends? The authors use 0.1 m/year for observed well-based trends, 0.01 m/year for raster trends, and 0.001 m/year for spatially aggregated trends. I would argue 0.1 m/year is not a negligible trend. Over 75 years a change of 7.5 m in a region with a shallow water table could either cause widespread groundwater flooding or cause traditional wells to dry up. Regardless, the authors should pick a threshold, justify it, and apply it consistently.
L 153: This assumption [] in many settings
L 382: Here you should state if the comparison is to the bias-corrected results or the raw GLOBGM predictions
L 413: If I understand correctly, the authors remove grid cells with negligible slope before taking the spatial average? What is the justification for this?
Based on Figure E1 it seems that trend magnitudes greater than 0.02 occur almost nowhere. Is this true? Or is the legend on Figure E1 incorrect? It seems inconsistent with Figure 9.
L507: I think Fig. E1 is the wrong cross-reference.
Figure D1: This bar graph is quite confusing. At first I read it as '47.1% of locations with |slope|<0.1 have rs<0'. I think what it is trying to say is 'of the locations with rs<0, 47.1% have |slope|<0.1.
Fig 11: I think these are anomalies from the 1970 starting point? A negative WTD does not make sense.
L625: typo
L716: I do not see agreement in the Gulf of Guinea or in Niger. GLOBGM predict negative changes in these regions, GRACE is strongly positive.
L721: "Despite these structural differences..." The authors have only identified one structural difference.
L752: It's a bit disappointing that the authors chose to compare their results only to the two example studies from the US that I pointed out. A quick google search shows that there are many more (large-scale) regional modelling studies that could be compared, eg. for Brazil, the UK, Germany, and China. Also, Meixner predicted decreased recharge across most of the southwestern US, and no change/slight increases in northern systems. Is this really in line with the GLOBGM v1.1 predictions of "rising water tables across most of the western contiguous United States"?
Lastly, there are a number of typos throughout the revised manuscript, and I would encourage the authors to proofread carefully. |
This study reports the results of a global groundwater model (GLOBGM v1.1) run at 1 km resolution under 3 CMIP6 scenarios and 5 climate models. Van Jaarsveld et al predict that groundwater levels will rise, on average, on most continents but that regions of groundwater depletion will persist. The study is ambitious in scope, and I applaud the authors for an attempt to make climate change predictions for an important component of the earth system. Unfortunately, the study suffers from several issues in methodology, validation, and interpretation, and because of this I believe the results are not reliable. Given the potential for the results to inform groundwater management policy, I think it would be inappropriate and potentially harmful to publish the manuscript. I unfortunately must recommend rejection.
My primary concern is the accuracy of the model and how it is evaluated. Overall, at approximately one third of wells the correlation to observations is negative: the model predicts the wrong direction of change. In my opinion, this is not good enough to make predictions about future trends. Based on Figure 5, panel a, the regions that show poor performance do not line up particularly well with the regions in Figure 4 where "GLOBGM has been shown to provide less accurate results". It seems the poor performance is primarily in places with high GW abstraction. Also, at 10%-20% of the wells the bias ratio is less than 0.1, meaning the simulated water table depth is more than 10 times too low? Again, this seems quite poor. Given that the model struggles to predict groundwater levels in regions where groundwater data exist and the subsurface is relatively well-understood, I am skeptical of the predictions for the rest of the world.
I looked at some previous publications related to this model, and similar concerns regarding accuracy have been raised (https://doi.org/10.5194/gmd-2022-226-RC1 , https://doi.org/10.5194/egusphere-2024-1025-RC3). In those cases the authors argued that the performance cannot be expected to be high, given that the model was not calibrated. They argued that the model is still valuable because "The philosophy of these models is that they try to capture the right processes and do not rely on calibration to correct for errors in process representation, parameterization and/or meteorological forcing data." Now the authors present a model that is calibrated, and wherein at least one of the inputs (groundwater recharge) in statistically downscaled to match observations. In addition, the authors use a machine learning model to adjust the predicted groundwater levels to match observations. So now the model is not conserving mass, and nevertheless the performance remains unsatisfying. I think it is time to either a) overhaul this model to try to achieve better global performance, b) focus only regions where the model matches historical observation, or c) return to a coarse-resolution model that might simulate regional dynamics without claiming 1-km resolution.
The use of a threshold of -0.41 was proposed by Knoben et al. (2019) for the KGE, but it is not valid for the KGE-NP. The threshold should be 0. Suppose the GW head varies around 50 m with a standard deviation of 5 m. For the mean of observations, I get a KGE-NP of -0.0008. This occurs because the Alpha component goes to 1 if distribution is close to symmetrical and not close to 0. I encourage the authors to verify this with their observational data.
The performance of the groundwater recharge downscaling (Section 2.1.2) is not reported. Was any cross-validation attempted? Moeck's (2020) data are not uniformly globally distributed. If you remove the Australian data from the training data, for example, can you predict those recharge rates with the regression model? In addition, limiting GWR_corrected to less than or equal to precipitation is not well-justified. In most cases this is probably a quite liberal constraint but in agricultural areas return flows can lead to recharge that exceeds precipitation. The choice of a multiplicative correction factor is also not justified.
The Machine Learing Bias Correction (2.4.3) should be also cross-validated regionally. The appendix indicates that the R-squared value is 0.6, which is already not very convincing as far as machine learning algorithms go. Can the model predict the bias for regions not included in training data? This is what would be required to apply this bias correction globally. Also, the labels in Figure A1 (a) are not defined.
My second major concern is the interpretation and discussion of results.
I would have expected to see more specific discussion of the future trends in different regions. What will happen to the major agricultural regions? Grouping by continents is not very informative. The fact that groundwater will rise in the Andes does not help northern Colombia and Central America, where groundwater levels could fall by ~50 m by the end of the century! Similarly for rising groundwater in Tibet and falling levels in the Ganges-Bramhmaputra. Also note that the majority of the rise seems to be happening in the Mountain regions where the authors say the model is less reliable. And what is causing the trend reversal in northwest North America?
A more robust discussion of uncertainty is needed, and the results should be compared to the literature. Given (a) the performance of the model in data-rich regions, (b) uncertain parameterization of the model in data-scarce regions, and (c) the uncertainty of the GCMs (you have included only five, and only one variant from each), how confident can we be in the projected trends? How do they compare to previous studies? For example, across the United States, Condon et al. (2020) predict deepening water tables across the US under warming. Meixner et al (2016) predicted decreased recharge over the southwestern US, little change of the northwest, and also that mountain recharge would decrease. Other regional assessments exist and should be compared as well.
I have some further minor comments and suggestions for the authors should they consider revising and resubmitting, in no particular order:
The authors state that forcing for the model are abstraction, recharge, and discharge (L227). Is this correct? If so, what is the model doing other than accounting for inputs and outputs? Why does the dynamic drainage elevation (L104) matter if groundwater discharge is already prescribed?
Equations 12-15: variables are not consistently labelled. What is alpha in equation 13? Is this different from the alpha in eq 10? Is WTDsim different from Wsim?
There is no confining layer over Canada at all? Surely this could be improved.
Two of the GCMs (IPSL-CM6A-LR and UKESM1-0-LL) in your ensemble are 'hot' - they have and Equilibrium Climate Sensitivity (ECS) and Transient Climate Response (TCR) above the assessed 'very likely' ranges estimated by the IPCC. This means that their projected warming is probably too pessimistic for a given scenario. The recommendation is, if the warming trajectory is important (which I think it is here) to use only models that lie within the likely range (Hausfather, 2022). Consider reporting the ensemble mean just for the three models that do lie within the 'likely' range, and including the others in an appendix.
The model struggles to simulate well-based GW levels and trends, particularly in places where GW use is high. Maybe that is not surprising, given that your water use data (Lange et al, 2021) is originally at 0.5 degree resolution, and the recharge is also based on downscaling coarser-resolution data. Perhaps simulating well-based trends is too difficult a task at the global scale. Does the model accurately simulate regional trends? I suggest the authors perform a simple experiment: aggregate the well-based data at increasingly coarse resolutions (say, 1 km, 2 km, 10 km, 50 km, and 100 km) and do the same for the model data, and then calculate your performance metrics at each resolution. I would expect you'll find performance will improve with coarser resolutions. Then focus on reporting results at the finest resolution that provides an adequate match to observations.
Eq. 17: Is this missing a month index?
Figure 6a: The color scale for this figure should be the same as for the uncorrected WTD bias (Figure 2).
Figure 7: It's unclear why these regions were chosen for insets. In any case they provide only about 2X magnification. I suggest removing the insets.
Discussion: The discussion of modelling choices, calibration, and advances is somewhat incongruous with ESD. I would expect this in a journal like Geoscientific Model Development but I think for ESD the discussion should be more general, and focus more on the implications.
L424 - The beginning of this paragraph seems to be missing. "[Major rivers], such as..."?
L435: Are these rising water tables in the north robust? Do they match observations? The authors state that 'This could conceivably be due to climate change enhancing precipitation and groundwater recharge dynamics'. It seems to me there is no need to speculate here - are those two processes actually occurring in the model?
Line 245: You could report the number of CPU-hours also. In addition, it would be responsible to report the CO2 footprint of these computations. This can be estimated as:
(node-hours) * (12 nodes/ # nodes at Snellius) * (power usage at Snellius) * (carbon intensity of Dutch grid)
Based on a quick search I get:
551 h * (12/1557) * (1200 kW) * 235 gCO2e/kWh = 1200 kg CO2e, about equal to a round trip flight ticket from Amsterdam to Beijing.
3.2.2 - Are these values for the ML-corrected data or the raw model output?
If the model purports to include anthropogenic influences on the water table, why are regions with anthropogenic influence excluded from calibration?
L451: I would not call this 'disagreement' - rather, divergence in scenarios, or you could say the direction of change is scenario-dependent.
L288 'The performance of the model' - which model? GLOBGM or the XGBoost model?
Table 4: it would be more useful to report the absolute bias, so negative and positive biases do not cancel each other out. Did the calibration reduce the mean absolute bias?
Section 3.1.1. This section seems more like methods than results.
L390 and L473: The shallower depths show better correlations for monthly data but not for yearly. That suggests that at shallow depths you can better capture water table variation related to precipitation seasonality, but not long-term variation and change. Seasonal variation is probably only strong for shallow wells, so it makes sense to me that the difference in performance disappaears in the yearly time series. I therefore disagree that "seasonal dynamics are more challenging to capture than inter-annual trends".
Figure 10: is the y-axis in m? What does the uncertainty represent - standard deviation of the five GCMS? The GSWP3-W5E5 simulation should also be plotted on these graphs for comparison.
References
Condon, L. E., Atchley, A. L., and Maxwell, R. M.: Evapotranspiration depletes groundwater under warming over the contiguous United States, Nat Commun, 11, 873, https://doi.org/10.1038/s41467-020-14688-0, 2020.
Hausfather, Z., Marvel, K., Schmidt, G. A., Nielsen-Gammon, J. W., and Zelinka, M.: Climate simulations: recognize the ‘hot model’ problem, Nature, 605, 26–29, https://doi.org/10.1038/d41586-022-01192-2, 2022.
Knoben, W. J. M., Freer, J. E., and Woods, R. A.: Technical note: Inherent benchmark or not? Comparing Nash-Sutcliffe and Kling-Gupta efficiency scores, Catchment hydrology/Modelling approaches, https://doi.org/10.5194/hess-2019-327, 2019.
Meixner, T., Manning, A. H., Stonestrom, D. A., Allen, D. M., Ajami, H., Blasch, K. W., Brookfield, A. E., Castro, C. L., Clark, J. F., Gochis, D. J., Flint, A. L., Neff, K. L., Niraula, R., Rodell, M., Scanlon, B. R., Singha, K., and Walvoord, M. A.: Implications of projected climate change for groundwater recharge in the western United States, Journal of Hydrology, 534, 124–138, https://doi.org/10.1016/j.jhydrol.2015.12.027, 2016.