<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">ESD</journal-id><journal-title-group>
    <journal-title>Earth System Dynamics</journal-title>
    <abbrev-journal-title abbrev-type="publisher">ESD</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Earth Syst. Dynam.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">2190-4987</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/esd-17-23-2026</article-id><title-group><article-title>Enhanced climate reproducibility testing with false discovery rate correction</article-title><alt-title>Reproducibility testing with FDR</alt-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Kelleher</surname><given-names>Michael E.</given-names></name>
          <email>kelleherme@ornl.gov</email>
        <ext-link>https://orcid.org/0000-0002-9286-2630</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Mahajan</surname><given-names>Salil</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>Computational Hydrology and Atmospheric Sciences Group, Oak Ridge National Laboratory, 1 Bethel Valley Rd, Oak Ridge TN, USA</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Michael E. Kelleher (kelleherme@ornl.gov)</corresp></author-notes><pub-date><day>6</day><month>January</month><year>2026</year></pub-date>
      
      <volume>17</volume>
      <issue>1</issue>
      <fpage>23</fpage><lpage>39</lpage>
      <history>
        <date date-type="received"><day>16</day><month>May</month><year>2025</year></date>
           <date date-type="rev-request"><day>28</day><month>May</month><year>2025</year></date>
           <date date-type="rev-recd"><day>24</day><month>October</month><year>2025</year></date>
           <date date-type="accepted"><day>21</day><month>November</month><year>2025</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2026 Michael E. Kelleher</copyright-statement>
        <copyright-year>2026</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://esd.copernicus.org/articles/17/23/2026/esd-17-23-2026.html">This article is available from https://esd.copernicus.org/articles/17/23/2026/esd-17-23-2026.html</self-uri><self-uri xlink:href="https://esd.copernicus.org/articles/17/23/2026/esd-17-23-2026.pdf">The full text article is available as a PDF file from https://esd.copernicus.org/articles/17/23/2026/esd-17-23-2026.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d2e89">Simulating the Earth's climate is an important and complex problem, thus climate models are similarly complex, comprised of millions of lines of code. In order to appropriately utilize the latest computational and software infrastructure advancements in Earth system models running on modern hybrid computing architectures to improve their performance, precision, accuracy, or all three; it is important to ensure that model simulations are repeatable and robust. This introduces the need for establishing statistical or non-bit-for-bit reproducibility, since bit-for-bit reproducibility may not always be achievable. Here, we propose a short-simulation ensemble-based test for an atmosphere model to evaluate the null hypothesis that modified model results are statistically equivalent to that of the original model. We implement this test in version 2 of the US Department of Energy's Energy Exascale Earth System Model (E3SM). The test evaluates a standard set of output variables across the two simulation ensembles and uses a false discovery rate correction to account for multiple testing. The false positive rates of the test are examined using re-sampling techniques on large simulation ensembles and are found to be lower than the currently implemented bootstrapping-based testing approach in E3SM. We also evaluate the statistical power of the test using perturbed simulation ensemble suites, each with a progressively larger magnitude of change to a tuning parameter. The new test is generally found to exhibit more statistical power than the current approach, being able to detect smaller changes in parameter values with higher confidence.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>Biological and Environmental Research</funding-source>
<award-id>E3SM</award-id>
</award-group>
</funding-group>
</article-meta>
  <notes notes-type="copyrightstatement">
  
      <p id="d2e99">This manuscript has been authored by UT-Battelle, LLC under Contract No. DE-AC05-00OR22725 with the U.S. Department of Energy. The United States Government retains and the publisher, by accepting the article for publication, acknowledges that the United States Government retains a non-exclusive, paid-up, irrevocable, worldwide license to publish or reproduce the published form of this manuscript, or allow others to do so, for United States Government purposes. The Department of Energy will provide public access to these results of federally sponsored research in accordance with the DOE Public Access Plan (<uri>http://energy.gov/downloads/doe-public-access-plan</uri>, last access: 12 December 2025).</p>
</notes></front>
<body>
      


<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d2e114">Thousands of scientists and engineers work tirelessly in efforts to better understand and model the Earth's changing climate. A large portion of this effort has come from the development of Earth system models at modeling centers around the globe, which seek to simulate the atmosphere, ocean, cryosphere, land surface, and chemistry, among other components of the Earth system. These models are comprised of many millions of lines of code and are enormously complex projects worked on by many individuals, so the need arises to verify that contributions to the model code do not have unintended effects on answers produced. Thorough testing of the output from these models is routinely conducted, assessing if the results are identical (bit-for-bit) or not. If the results are not bit-for-bit identical, statistical checks are also conducted to ensure the simulated climate of the model or component has not significantly changed, unless that is the intended effect. A variety of methods are available <xref ref-type="bibr" rid="bib1.bibx28 bib1.bibx37 bib1.bibx2 bib1.bibx21 bib1.bibx23" id="paren.1"/>, each of which performs a statistical comparison between a reference ensemble and test simulation or ensembles. Here, an ensemble indicates a set of model runs each initialized with similar but slightly perturbed initial conditions, generally only at machine-precision levels.</p>
      <p id="d2e120">The Energy Exascale Earth System Model (E3SM, <xref ref-type="bibr" rid="bib1.bibx13 bib1.bibx10" id="altparen.2"/>) uses a suite of tests which run the model under a variety of configurations and methods using the Common Infrastructure for Modeling the Earth (CIME) software to setup, build, run, and analyze the model. The tests are run at varying frequencies from nightly to weekly, testing both the latest science and performance updates and maintenance branches, which only receive compatibility updates and should be bit-for-bit identical. A subset of these tests are statistical reproducibility tests (non-bit-for-bit tests), which determine through a comparison of control and perturbed ensembles, whether or not the simulated climate has changed as a result of modifications to the model code or infrastructure. These are the Time Step Convergence test (TSC, <xref ref-type="bibr" rid="bib1.bibx37" id="altparen.3"/>), the Perturbation Growth New test (PGN, related to the work in <xref ref-type="bibr" rid="bib1.bibx33" id="altparen.4"/>), the multi-testing Kolmogorov-Smirnov test (MVK, <xref ref-type="bibr" rid="bib1.bibx23" id="altparen.5"/>), and the MVK-Ocean test (MVK-O, <xref ref-type="bibr" rid="bib1.bibx20" id="altparen.6"/>). The first three test the reproducibility of the E3SM Atmosphere Model, (EAM), and the last tests the Model for Prediction Across Scales-Ocean (MPAS-O), the ocean component in E3SM.</p>
      <p id="d2e138">The TSC test evaluates numerical convergence by comparing ensemble differences at two small time step sizes (i.e., 1 and 2 s) using a Student's <inline-formula><mml:math id="M1" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> test on root mean squared difference (RMSD) values, under the assumption that numerical solutions should converge as the time step decreases <xref ref-type="bibr" rid="bib1.bibx37" id="paren.7"/>. The PGN test, in contrast, assesses stability by comparing the state of the atmosphere after one time step across perturbed ensemble members, identifying whether small initial differences grow inconsistently through individual physics parameterizations. Unlike the shorter duration TSC and PGN tests, the MVK and MVK-O tests use year-long and two-year long simulation ensembles respectively, allowing them to assess the cumulative impact of code or configuration changes on the model's climatology after internal variability has saturated, providing a more robust evaluation of long-term climate statistics.</p>
      <p id="d2e151">MVK evaluates the null hypothesis that a modified model simulation ensemble is statistically equivalent to a baseline ensemble. It applies a two-sample Kolmogorov-Smirnov test to over 100 output variables and counts how many show statistically significant differences between the two ensembles. If this count exceeds a critical value threshold as expected from internal variability and derived via bootstrapping, the two simulations are considered to have different simulated climates <xref ref-type="bibr" rid="bib1.bibx21" id="paren.8"/>. Further details on the MVK are provided in Sect. <xref ref-type="sec" rid="Ch1.S2.SS1"/>.</p>
      <p id="d2e160">Here, we propose a new testing approach that improves on the MVK and provides a two-fold benefit over it. Firstly, from a human usability perspective, it is desirable for the test to reduce the number of false positives (Type I error rates; where two simulations are erroneously labeled as statistically different), without affecting the false negative rates (Type II error rates; where two simulation ensembles are erroneously labeled as statistically similar). Operationally, MVK has been exhibiting a false positive rate of about 7.5 %, despite the prescribed significance level of 5 %,  since its induction into the test suite (as discussed in Sect. <xref ref-type="sec" rid="Ch1.S4.SS7"/>). Other previous works suggest that implementing a false discovery rate (FDR) correction (by adjusting significance thresholds based on the number and rank of <inline-formula><mml:math id="M2" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> values, discussed in more detail in Sect. <xref ref-type="sec" rid="Ch1.S2.SS3"/>), reduces the false positive rate when multi-testing as compared to using bootstrapping-derived critical value threshold for hypothesis testing <xref ref-type="bibr" rid="bib1.bibx35 bib1.bibx41 bib1.bibx42 bib1.bibx20" id="paren.9"/>. This is the methodology used by MVK. Some frameworks, including those presented in <xref ref-type="bibr" rid="bib1.bibx20" id="text.10"/>, <xref ref-type="bibr" rid="bib1.bibx43" id="text.11"/>, seek to reduce false positive rates, and do so successfully, but at a greater computational cost, as they work with grid-box level data. The work in <xref ref-type="bibr" rid="bib1.bibx43" id="text.12"/> compares larger sub-ensemble sizes which reduces the false positive rate lower than an FDR correction, but adds to the computational cost, while <xref ref-type="bibr" rid="bib1.bibx20" id="text.13"/> uses the same ensemble sizes, but applies an FDR correction at each grid box in the ocean model. The new testing approach for the atmosphere model thus implements an FDR correction when evaluating multiple variables in the atmosphere model output. Secondly, it will be useful to reduce the computational cost and time of deriving the critical value thresholds for the MVK. The bootstrapping procedure is time-consuming and computationally expensive because of its need for large control ensembles. And, it would need to be conducted again after a significant enough departure from the original model code to establish a new critical value threshold. Substantial model code changes to numerics and physics can alter internal variability and shift the distributional properties of output fields. FDR correction theoretically asserts the critical value threshold for global null hypothesis evaluation (as discussed in Sect. <xref ref-type="sec" rid="Ch1.S2.SS3"/>) is 1 or more rejected field, thus eliminating the need for conducting a large control ensemble simulation to determine this threshold. A large control ensemble is required, though, to demonstrate that the theoretical critical value is equivalent or better than the one arrived at by bootstrap sampling of the large ensemble.</p>
      <p id="d2e192">This paper thus seeks to answer two questions: can a testing framework using FDR correction reduce the number of false positives in an operational setting as compared to the MVK without impacting its statistical power (false negative rates) and can it eliminate the need for extensive and expensive analysis of large ensembles each time the model is updated significantly? We also explore the impact of applying the BH-FDR correction to different underlying statistical test used to evaluate the differences in distribution, using the Mann-Whitney U and Cramér-von Mises tests in addition to the K-S test.</p>
      <p id="d2e195">The following section discusses the MVK test and its pitfalls in more detail and describes the FDR approach as implemented here. Section <xref ref-type="sec" rid="Ch1.S3"/> lists the simulation ensembles conducted to evaluate both the MVK and the new testing framework. Section <xref ref-type="sec" rid="Ch1.S4"/> discusses our results on the evaluation of the false positive and negative rates of these testing frameworks. Finally, our results are summarized in Sect. <xref ref-type="sec" rid="Ch1.S5"/> with a brief discussion of caveats of the study and future direction.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Methods</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Multi-testing Kolmogorov-Smirnov Test</title>
      <p id="d2e219">The multi-testing Kolmogorov-Smirnov (MVK) test, as implemented in the CIME for the atmosphere model of E3SM, compares two independent <inline-formula><mml:math id="M3" display="inline"><mml:mrow><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">30</mml:mn></mml:mrow></mml:math></inline-formula> member ensembles. It is used in nightly testing, comparing a baseline ensemble, generated after each approved “climate changing” code modification, and a test ensemble, newly generated each day. The baseline ensemble is generated following approvals from domain scientists who have expertise related to the newly introduced code. The E3SM model is run at “ultra-low” resolution, (called ne4pg2, <inline-formula><mml:math id="M4" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">7.5</mml:mn></mml:mrow></mml:math></inline-formula>° atmosphere), for 14 months. The first two months are discarded as the system reaches quasi-equilibrium. Annual global means of each of the 120 standard output fields of the E3SM Atmosphere Model (EAM) are then computed for each ensemble member. Then, for each field, the null hypothesis (<inline-formula><mml:math id="M5" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, also referred to as the local null hypothesis here) is evaluated. The local null hypothesis asserts that the sample distribution function of the annual global mean of that field estimated from the baseline ensemble (from <inline-formula><mml:math id="M6" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> data points) is statistically similar to that of the new ensemble is evaluated. The two sample Kolmogorov-Smirnov (K-S) test at a significance level of <inline-formula><mml:math id="M7" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula> is used for testing <inline-formula><mml:math id="M8" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>. The K-S test is a non-parametric statistical test to compare cumulative or empirical distribution functions. The larger null hypothesis (<inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, also referred to as the global null hypothesis here) that the two ensembles have identical simulated climates is then evaluated for a significance level of <inline-formula><mml:math id="M10" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula>. The test statistic, <inline-formula><mml:math id="M11" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, for testing this larger null hypothesis is defined as the number of fields that reject <inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>. <inline-formula><mml:math id="M13" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> is rejected if the number of fields rejecting <inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> is greater than a critical value threshold (found to be 13, <xref ref-type="bibr" rid="bib1.bibx21" id="altparen.14"/>), and the test issues a “fail”. The null distribution of <inline-formula><mml:math id="M15" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, and hence the critical value threshold at a significance level of <inline-formula><mml:math id="M16" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula>, was empirically derived by randomly sampling two <inline-formula><mml:math id="M17" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula>-member ensembles from a 150-member control ensemble of an earlier version of E3SM, and computing <inline-formula><mml:math id="M18" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, 500 times <xref ref-type="bibr" rid="bib1.bibx21" id="paren.15"/>. While regional means, local diagnostics, and extremes can provide valuable perspectives on ensemble behavior, here we focus on annual global means, following previous work <xref ref-type="bibr" rid="bib1.bibx2 bib1.bibx21" id="paren.16"/>. Atmospheric variability is highly heterogeneous in both space and time, with sharp gradients and transient fluctuations that introduce substantial sampling noise into regional ensemble diagnostics. Global averaging suppresses stochastic, weather-driven variability and reduces dimensionality, enabling systematic differences between two ensembles to be more clearly identified. Further, extremes require much larger ensembles to be reliably assessed, as <xref ref-type="bibr" rid="bib1.bibx21" id="text.17"/> demonstrated. With ensembles of about sixty members, extremes of temperature and precipitation were often statistically indistinguishable even when mean distributions diverged, underscoring their limited sensitivity in this context.</p>
<sec id="Ch1.S2.SS1.SSSx1" specific-use="unnumbered">
  <title>Potential pitfalls of multi-testing K-S Test</title>
      <p id="d2e391">MVK requires a large control ensemble in order to capture the variability of the model and establish proper critical value thresholds for the number of rejected local null hypotheses before rejection of the global null hypothesis is considered <xref ref-type="bibr" rid="bib1.bibx41" id="paren.18"/>. It also, as discussed in <xref ref-type="bibr" rid="bib1.bibx41" id="text.19"/>, does not put weight onto fields rejected very strongly. That is, those fields with extremely small <inline-formula><mml:math id="M19" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> values. If a few fields are rejected with near certainty, this likely indicates global significance. But, if the total number of these fields does not exceed the predetermined threshold, the overall result is global null hypothesis acceptance, which results in a lower power to detect differences. Finally, in practice this approach has a larger Type I (false positive) error (see Sect.  <xref ref-type="sec" rid="Ch1.S4.SS2"/>, <xref ref-type="sec" rid="Ch1.S4.SS7"/>; Table <xref ref-type="table" rid="T2"/>) than desired. While possible to decrease the significance level <inline-formula><mml:math id="M20" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> to reduce the false positive rates, this in turn decreases the statistical power of the test to detect small changes which is not desirable.</p>
</sec>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Additional statistical tests</title>
      <p id="d2e430">The Kolmogorov-Smirnov test used in the MVK framework is just one non-parametric test available to ascertain differences between distributions of random variables. To add robustness to the assessment of distributional differences, the Mann-Whitney <inline-formula><mml:math id="M21" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> (also called the Wilcoxon rank sum) and Cramér-von Mises tests were also performed to compare each set of ensembles. The Mann-Whitney <inline-formula><mml:math id="M22" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> test (M-W hereafter), ranks all samples from the two groups, and compares the sum of those ranks between the two groups <xref ref-type="bibr" rid="bib1.bibx25" id="paren.20"/>. The Cramér-von Mises test (C-VM hereafter) used here compares two empirical distributions using the quadratic distance between them <xref ref-type="bibr" rid="bib1.bibx1" id="paren.21"/>. These tests present an advantage in statistical power over the K-S test, as more weight is given to the distribution tails <xref ref-type="bibr" rid="bib1.bibx12" id="paren.22"/>, resulting in higher sensitivity to extreme values, though at the cost of being more sensitive to outliers. Both additional tests are also non-parametric tests, meaning that limited assumptions are asserted about the shape of the data.</p>
</sec>
<sec id="Ch1.S2.SS3">
  <label>2.3</label><title>False Discovery Rate Correction</title>
      <p id="d2e464">When conducting simultaneous multiple null hypothesis tests as in the MVK above, Type I error rate inflation occurs, leading to a higher overall probability of false positives, since the number of expected null hypothesis rejections increases with each additional test <xref ref-type="bibr" rid="bib1.bibx3" id="paren.23"/>. This means that increased false positives can occur when testing many <inline-formula><mml:math id="M23" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> values at once. While a significance level (e.g., <inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula>) controls the probability of a false positive for a single test, conducting many tests increases the chance that at least one result will appear significant purely by chance even if all the null hypotheses are true. This implies that the overall proportion of false discoveries can become unacceptably high, undermining the reliability of the findings. MVK accounts for multi-testing by using a re-sampling strategy, which was found to give similar results as compared to permutation testing <xref ref-type="bibr" rid="bib1.bibx23" id="paren.24"/>. In addition to permutation testing and bootstrapping approaches, other approaches for correcting Type I error rate inflation associated with multi-testing include family wise error rate correction (e.g. Bonferroni correction), and the false discovery rate (FDR) <xref ref-type="bibr" rid="bib1.bibx41 bib1.bibx35" id="paren.25"/>, the latter of which is used here. The Bonferroni correction adjusts the significance level by dividing the desired significance level, <inline-formula><mml:math id="M25" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula>, by the number of tests conducted, but is known to have reduced statistical power <xref ref-type="bibr" rid="bib1.bibx41 bib1.bibx35" id="paren.26"/>. Here, we use the Benjamini-Hochberg (BH) FDR correction approach <xref ref-type="bibr" rid="bib1.bibx3" id="paren.27"/>. The BH-FDR approach has been shown to effectively control for Type I error rate inflation while also exhibiting more power than other approaches <xref ref-type="bibr" rid="bib1.bibx35 bib1.bibx41" id="paren.28"/>. It is widely used in Earth system studies for spatial analysis <xref ref-type="bibr" rid="bib1.bibx41 bib1.bibx32 bib1.bibx40" id="paren.29"/> and has also been applied to solution reproducibility testing of ocean models <xref ref-type="bibr" rid="bib1.bibx20" id="paren.30"/>. Other methods of false discovery rate correction (<xref ref-type="bibr" rid="bib1.bibx4" id="altparen.31"/>, Bonferroni adjustment) were also examined and found to have lower power than the BH-FDR method used here. These corrections remove more rejections than the BH-FDR correction and result in fewer global null hypothesis rejections, so they are less sensitive to small parameter changes. While the original BH-FDR <xref ref-type="bibr" rid="bib1.bibx3" id="paren.32"/> description asserted independence, it has been found in more recent work (e.g. <xref ref-type="bibr" rid="bib1.bibx4 bib1.bibx35" id="altparen.33"/>) that the independence of the <inline-formula><mml:math id="M26" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> values is not a strict requirement.</p>
      <p id="d2e535">Similar to MVK, we use the two sample K-S, the two sample C-VM, and the two sample M-W tests to evaluate the local null hypothesis (<inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) for each field that their sample distribution functions are statistically identical across the two ensembles. Also, similar to the MVK, the global null hypothesis (<inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) being tested is that the two ensembles are statistically similar. We use the <monospace>statsmodels</monospace> <xref ref-type="bibr" rid="bib1.bibx34" id="paren.34"/> Python package for applying BH to reproducibility testing. This corrects the critical threshold, <inline-formula><mml:math id="M29" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula>, for evaluating local null hypothesis (<inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) rejection with the <inline-formula><mml:math id="M31" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>th sorted <inline-formula><mml:math id="M32" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> value, where <inline-formula><mml:math id="M33" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> is defined in Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>) (see Eq. 2 of <xref ref-type="bibr" rid="bib1.bibx35" id="altparen.35"/>, and Eq. 3 of <xref ref-type="bibr" rid="bib1.bibx42" id="altparen.36"/>).

            <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M34" display="block"><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">max⁡</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mi mathvariant="normal">…</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:munder><mml:mfenced open="[" close="]"><mml:mrow><mml:mi>i</mml:mi><mml:mo>:</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mo>≤</mml:mo><mml:msup><mml:mi>q</mml:mi><mml:mo>*</mml:mo></mml:msup><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced></mml:mrow></mml:math></disp-formula>

          Where <inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the <inline-formula><mml:math id="M36" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>th smallest <inline-formula><mml:math id="M37" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> value, <inline-formula><mml:math id="M38" display="inline"><mml:mi>m</mml:mi></mml:math></inline-formula> is the total number of <inline-formula><mml:math id="M39" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> values (equivalent to the number of hypothesis tests), and <inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:msup><mml:mi>q</mml:mi><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> is the chosen limit on the false discovery rate, typically chosen to be <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:msup><mml:mi>q</mml:mi><mml:mo>*</mml:mo></mml:msup><mml:mo>=</mml:mo><mml:mi mathvariant="italic">α</mml:mi></mml:mrow></mml:math></inline-formula>. This <inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:msup><mml:mi>q</mml:mi><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> is defined as the upper limit of the false positive rate using the FDR-BH methods, and is chosen at the 5 % level as it balances false negative and false positive rates. A decrease in <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:msup><mml:mi>q</mml:mi><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> would decrease false positives, but it would be offset by an increase in false negative errors. The result of Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>), <inline-formula><mml:math id="M44" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>, is thus the index of the largest <inline-formula><mml:math id="M45" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> value which is less than <inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:msup><mml:mi>q</mml:mi><mml:mo>*</mml:mo></mml:msup><mml:mi>i</mml:mi><mml:mo>/</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:math></inline-formula>, and all <inline-formula><mml:math id="M47" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> values with an index less than <inline-formula><mml:math id="M48" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> (e.g. <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:msub><mml:mi mathvariant="normal">…</mml:mi><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>) are rejected. Now, the global null hypothesis, <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, that can be framed as all <inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> are true, is rejected at the global significance level of <inline-formula><mml:math id="M52" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> if <italic>any</italic> <inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:msubsup><mml:mi>H</mml:mi><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:math></inline-formula> is rejected <xref ref-type="bibr" rid="bib1.bibx32 bib1.bibx41 bib1.bibx35" id="paren.37"/>. Thus, the global null hypothesis is rejected if <italic>any</italic> <inline-formula><mml:math id="M54" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> value is rejected, said another way, if <inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:mi>k</mml:mi><mml:mo>≥</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>. In other words, the critical value threshold for the test statistic, <inline-formula><mml:math id="M56" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, which is the number of fields rejecting <inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, is equal to one for <inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> for BH-FDR. This theoretical critical value threshold of one, when multi-testing with FDR corrections, has been used widely for field significance testing to address the multiple comparisons problem inherent in spatial analyses in the climate and meteorological studies <xref ref-type="bibr" rid="bib1.bibx41 bib1.bibx32 bib1.bibx6" id="paren.38"/>. More details about the BH-FDR as applied to earth system science can also be found in these studies.</p>
      <p id="d2e947">This BH-FDR as applied to reproducibility testing will be deemed useful here if it performs at least as well at detecting altered simulated climates (statistical power) as the MVK while also reducing the number of false positives. We investigate these characteristics for each approach using suites of simulation ensembles with controlled modifications in Sect. <xref ref-type="sec" rid="Ch1.S4"/>.</p>
</sec>
<sec id="Ch1.S2.SS4">
  <label>2.4</label><title>Illustration of FDR approach with test cases</title>
      <p id="d2e960">Figure <xref ref-type="fig" rid="F1"/> illustrates the FDR approach using two example simulation ensemble comparisons, described later in Sect. <xref ref-type="sec" rid="Ch1.S3"/>. At left, a control comparison is shown where each ensemble member differs by a machine precision perturbation only (expected to globally pass, but have local rejections), and at right a perturbed ensemble experiment is compared to the control with known solution changes (expected to be globally rejected). Note that several <inline-formula><mml:math id="M59" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> values in the control self comparison are below the threshold <inline-formula><mml:math id="M60" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula>, but none fall below <inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">α</mml:mi><mml:mi mathvariant="normal">FDR</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula>, leading to an acceptance of the global null hypothesis for this experiment. In the comparison between control and perturbed ensembles there are fewer <inline-formula><mml:math id="M62" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> values which fall below the corrected <inline-formula><mml:math id="M63" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">α</mml:mi><mml:mi mathvariant="normal">FDR</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> than fall below the nominal <inline-formula><mml:math id="M64" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula>, despite this, the global null hypothesis is still rejected, as at least one field is rejected. This is similar to Fig. 2 from <xref ref-type="bibr" rid="bib1.bibx35" id="text.39"/>, which also illustrates the impact of a variable <inline-formula><mml:math id="M65" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula>.</p>

      <fig id="F1" specific-use="star"><label>Figure 1</label><caption><p id="d2e1030">Sorted <inline-formula><mml:math id="M66" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> values for two selected experiments, shown in blue, orange, and green for the Kolmogorov-Smirnov, Cramér-von Mises, and Mann-Whitney <inline-formula><mml:math id="M67" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> tests respectively. These lines are thicker where <inline-formula><mml:math id="M68" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> values are rejected based on BH-FDR correction. The dashed grey line shows the nominal <inline-formula><mml:math id="M69" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula>, and the dashed red line shows the FDR corrected <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:msup><mml:mi mathvariant="italic">α</mml:mi><mml:mi mathvariant="normal">FDR</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mi>q</mml:mi><mml:mo>*</mml:mo></mml:msup><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi>i</mml:mi><mml:mo>/</mml:mo><mml:mi>m</mml:mi></mml:mrow></mml:math></inline-formula>.</p></caption>
          <graphic xlink:href="https://esd.copernicus.org/articles/17/23/2026/esd-17-23-2026-f01.png"/>

        </fig>

</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Simulation Ensembles</title>
      <p id="d2e1106">We conduct a suite of multi-member ensembles to evaluate the false positive and false negative rates of BH-FDR approach. Additionally, having undergone significant updates to software <xref ref-type="bibr" rid="bib1.bibx13" id="paren.40"/>, we reassess the critical value threshold of MVK at which two ensembles of E3SMv2.1 are statistically distinguishable (the critical value threshold having been previously computed from a control ensemble of an earlier model version) using this ensemble suite. Each ensemble is generated by varying the initial conditions by near-machine precision perturbations (<inline-formula><mml:math id="M71" display="inline"><mml:mrow><mml:mi>O</mml:mi><mml:mo>(</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>)) to the temperature field at each grid point for each ensemble member. One of these ensembles is the control ensemble which refers to the unmodified default E3SMv2.1 model with default values of all tuning parameters. Other generated ensembles are differentiated from the control ensemble by varying the value of a tuning parameter from its default value by different magnitudes (see Table <xref ref-type="table" rid="T1"/> for details). Here, we vary three tuning parameters separately. Version 1 of E3SM was found to be highly sensitive to a parameter termed <monospace>clubb_c1</monospace>, somewhat sensitive to <monospace>zmconv_c0_ocn</monospace>, and weakly sensitive to the <monospace>effgw_oro</monospace> <xref ref-type="bibr" rid="bib1.bibx30" id="paren.41"/>, similar parametric sensitivity was found in the Community Atmosphere Model version 6 (CAMv6) <xref ref-type="bibr" rid="bib1.bibx11" id="paren.42"/>. <monospace>clubb_c1</monospace> is the constant associated with dissipation of variance of <inline-formula><mml:math id="M72" display="inline"><mml:mover accent="true"><mml:mrow><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mo>′</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:math></inline-formula> (where <inline-formula><mml:math id="M73" display="inline"><mml:mi>w</mml:mi></mml:math></inline-formula> is vertical wind speed), <monospace>zmconv_c0_ocn</monospace> is the deep convection precipitation efficiency over ocean grid-points, and <monospace>effgw_oro</monospace> is the gravity wave drag intensity (see Table 1 of <xref ref-type="bibr" rid="bib1.bibx30" id="altparen.43"/>). We choose these three parameters to generate ensembles to capture a range of sensitivities in version 2 of E3SM. As in the operational case described in Sect. <xref ref-type="sec" rid="Ch1.S2.SS1"/>, the simulation ensembles are run at “ultra-low” resolution with a 14-month simulation duration. A 120-member ensemble is generated for each tuning parameter change, along with the control ensemble.</p>

<table-wrap id="T1"><label>Table 1</label><caption><p id="d2e1190">List of simulations for each tuning parameter. Bold font indicates value in control simulation.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Parameter</oasis:entry>
         <oasis:entry colname="col2">% Change</oasis:entry>
         <oasis:entry colname="col3">Parameter value</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>effgw_oro</monospace></oasis:entry>
         <oasis:entry colname="col2"><bold>0.0</bold></oasis:entry>
         <oasis:entry colname="col3"><bold>0.375</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">1.0</oasis:entry>
         <oasis:entry colname="col3">0.3788</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">10.0</oasis:entry>
         <oasis:entry colname="col3">0.4125</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">20.0</oasis:entry>
         <oasis:entry colname="col3">0.4500</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">30.0</oasis:entry>
         <oasis:entry colname="col3">0.4875</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">40.0</oasis:entry>
         <oasis:entry colname="col3">0.5250</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">50.0</oasis:entry>
         <oasis:entry colname="col3">0.5625</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>clubb_c1</monospace></oasis:entry>
         <oasis:entry colname="col2"><bold>0.0</bold></oasis:entry>
         <oasis:entry colname="col3"><bold>2.400</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">1.0</oasis:entry>
         <oasis:entry colname="col3">2.424</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">3.0</oasis:entry>
         <oasis:entry colname="col3">2.472</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">5.0</oasis:entry>
         <oasis:entry colname="col3">2.520</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">10.0</oasis:entry>
         <oasis:entry colname="col3">2.640</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><monospace>zmconv_c0</monospace></oasis:entry>
         <oasis:entry colname="col2"><bold>0.0</bold></oasis:entry>
         <oasis:entry colname="col3"><bold>0.0020</bold></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">0.5</oasis:entry>
         <oasis:entry colname="col3">0.00201</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">1.0</oasis:entry>
         <oasis:entry colname="col3">0.00202</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">3.0</oasis:entry>
         <oasis:entry colname="col3">0.00206</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">5.0</oasis:entry>
         <oasis:entry colname="col3">0.00210</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e1423">The <monospace>effgw_oro</monospace> and <monospace>clubb_c1</monospace> ensembles were run on Argonne National Laboratory's Chrysalis machine and built using the Intel compiler v20.0.4, while the <monospace>zmconv_c0_ocn</monospace> ensembles were run on US Air Force HPC11 at Oak Ridge National Laboratory's Oak Ridge Leadership Computing Facility (OLCF), compiled using the GNU compilers v13.2. An additional control ensemble was also performed on HPC11.</p>
      <p id="d2e1436">To further evaluate the BH-FDR approach, we conduct two additional ensemble simulations on Chrysalis to test model sensitivity to compiler optimization choices. The default optimization flag for E3SM is “-O3”, and is used to generate the control and other ensembles with tuning parameter changes. The two different optimization test ensembles are titled <monospace>opt-O1</monospace> and <monospace>fastmath</monospace>, and are compiled with optimization flags “-O1” and “-O3 -fp-model=fast” respectively. Previous work suggests that using optimization “-O1” is expected to not produce a significantly different simulated climate than the default <xref ref-type="bibr" rid="bib1.bibx2 bib1.bibx21" id="paren.44"/>. <xref ref-type="bibr" rid="bib1.bibx21" id="text.45"/> found that using optimization flag “fast” with “Mvect” resulted in a statistically different climate compared to the default, which used an optimization of “-O2” using the PGI compiler.</p>
      <p id="d2e1451">The “-O1” optimization turns off most of the aggressive optimizations used for the “-O3” level, including loop vectorization, loop unrolling, and global register allocation, while enabling the “-fp-model=fast” fast floating point model allows the compiler to be less strict in its handling of floating point arithmetic <xref ref-type="bibr" rid="bib1.bibx17" id="paren.46"/>. This means using “-O1”  in place of “-O3” ought to result in slower operation but similar results, while using “-fp-model=fast” could result in different answers under specific conditions, including if there are “NaN” or not-a-number values present. Though in the case of E3SM, “-fp-model=fast” is already used in the compilation of several source files, thus adding it as a global option only changes those where it is not in use already.</p>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Results</title>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Estimating Critical Value Threshold for MVK</title>
      <p id="d2e1473">E3SMv2.1 has undergone several scientific feature changes as well as software infrastructure changes since the release of E3SMv0 which was used to estimate the critical value threshold for null hypothesis testing using MVK <xref ref-type="bibr" rid="bib1.bibx13" id="paren.47"/>. Here, we estimate the null distribution of the test statistic of MVK which is more representative of E3SMv2.1. We use a bootstrapping (re-sampling) strategy to derive the null distribution and the empirical critical value threshold of the test statistic, <inline-formula><mml:math id="M74" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, which is the number of variables falsely rejecting the true null hypothesis <inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> at <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula>, for the global null hypothesis using for the MVK test for E3SMv2.1. For this analysis, two 30 member ensembles were drawn, without replacement, from the 120 member control ensemble (with <inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mn mathvariant="normal">18</mml:mn></mml:msup></mml:mrow></mml:math></inline-formula> possible ways of drawing a 30 member ensemble). By drawing two ensembles from the same population, this establishes an expected value for how many fields would be rejected by random chance when using two ensembles which have the same simulated climate. <inline-formula><mml:math id="M78" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> is computed for each such draws of 30-member ensemble pairs and then this procedure is repeated 1000 times. This procedure was then applied to each parameter adjustment ensemble separately, comparing two random draws from each large ensemble, in an effort to expand the sample size to estimate the null distribution. It is also applied to the ensembles with changes to compiler optimizations, thus the empirical threshold for rejecting the global null hypothesis is computed from 2040 ensemble members. The null distribution of <inline-formula><mml:math id="M79" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> is representative of the internal variability of the model as the drawn ensemble pairs are part of the same population with ensemble members differing only in the initial conditions at machine precision level perturbations. Figure <xref ref-type="fig" rid="F2"/> illustrates the null distribution of <inline-formula><mml:math id="M80" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> with a box and whiskers plot derived from each 120-member ensemble using the K-S test. The critical value threshold is also estimated as the 95th percentile of <inline-formula><mml:math id="M81" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, which ranges from 10 to 13 for different ensembles.</p>

      <fig id="F2" specific-use="star"><label>Figure 2</label><caption><p id="d2e1555">Box plot of number of rejected output fields using the K-S test for each ensemble self-comparison. The solid center line of each box represents the median, the box represents the inter-quartile range (25th–75th percentiles), the whiskers are plotted at the 5th–95th percentiles, and outliers marked beyond that range. The dashed vertical line is the median of all 95th percentiles which is 11 fields.</p></caption>
          <graphic xlink:href="https://esd.copernicus.org/articles/17/23/2026/esd-17-23-2026-f02.png"/>

        </fig>

      <p id="d2e1564">We thus set the critical value threshold for MVK (K-S) at <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula> as the median of those values, found to be 11. This procedure was repeated using the other two statistical tests with similar results, though the thresholds for C-VM and M-W were both found to be larger, at 16.</p>
</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>False positive rates: Uncorrected and BH-FDR approach</title>
      <p id="d2e1587">The bootstrapping method described above in Sect. <xref ref-type="sec" rid="Ch1.S4.SS1"/> to determine the critical value threshold for global null hypothesis testing for MVK can also be used to estimate the false positive rates. Since ensemble pairs are drawn from the same population, each drawn pair that rejects the null hypothesis is a false positive. The thresholds for global rejection are set at K-S: 11, C-VM: 16, and M-W: 16 for null hypothesis testing at <inline-formula><mml:math id="M83" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula>. For the BH-FDR approach applied to each statistical test, any field rejecting the null hypothesis after false discovery rate correction implies a rejection of the larger null hypothesis.</p>
      <p id="d2e1604">For each of the thirteen 120-member ensembles conducted (control, tuning parameter changes and optimization change ensembles), a 1000 iteration bootstrapping analysis is performed separately, and the false positive rates are computed for each under each statistical test and the BH-FDR testing approach. Table <xref ref-type="table" rid="T2"/> details the false positive rates derived from each analysis. The mean of these 17 values for the MVK is 0.046, which can be expected to be near the prescribed <inline-formula><mml:math id="M84" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> since the critical value threshold was estimated from the same population (set of ensembles), although with different bootstrap samples, similarly for M-W and C-VM tests, the mean false positive rate are 0.048 and 0.049 respectively. The mean false positive rate under the BH-FDR approach is lower for all tests at 0.032, 0.035, and 0.040 for MVK, M-W, and C-VM respectively. Also, the 95th percentile of false positive rates for the 17 ensemble comparisons are also reduced BH-FDR approach for all statistical tests. The above indicate that the application of the FDR correction works as intended and based on the Lemma of Theorem 1 in <xref ref-type="bibr" rid="bib1.bibx3" id="text.48"/>, the FDR puts an upper limit to the level of false discovery at <inline-formula><mml:math id="M85" display="inline"><mml:mrow><mml:msup><mml:mi>q</mml:mi><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:math></inline-formula> (which here is chosen as <inline-formula><mml:math id="M86" display="inline"><mml:mrow><mml:msup><mml:mi>q</mml:mi><mml:mo>*</mml:mo></mml:msup><mml:mo>=</mml:mo><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula>), thus false positive rates lower than <inline-formula><mml:math id="M87" display="inline"><mml:mi mathvariant="italic">α</mml:mi></mml:math></inline-formula> are expected.</p>

<table-wrap id="T2" specific-use="star"><label>Table 2</label><caption><p id="d2e1660">False positive rates for self-comparison bootstraps (per 1000 iterations for each ensemble) at <inline-formula><mml:math id="M88" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula>, and the mean and 95th percentile false-positive rate  over all self-comparison bootstraps.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right" colsep="1"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right" colsep="1"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Statistical Test</oasis:entry>
         <oasis:entry namest="col2" nameend="col3" align="center">Kolmogorov-Smirnov (K-S) </oasis:entry>
         <oasis:entry namest="col4" nameend="col5" align="center" colsep="1">Mann-Whitney (M-W) </oasis:entry>
         <oasis:entry namest="col6" nameend="col7" align="center">Cramér-von Mises (C-VM) </oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Method</oasis:entry>
         <oasis:entry colname="col2">Uncorrected</oasis:entry>
         <oasis:entry colname="col3">BH-FDR Corrected</oasis:entry>
         <oasis:entry colname="col4">Uncorrected</oasis:entry>
         <oasis:entry colname="col5">BH-FDR Corrected</oasis:entry>
         <oasis:entry colname="col6">Uncorrected</oasis:entry>
         <oasis:entry colname="col7">BH-FDR Corrected</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col7">Simulation Ensemble </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Control</oasis:entry>
         <oasis:entry colname="col2">0.049</oasis:entry>
         <oasis:entry colname="col3">0.028</oasis:entry>
         <oasis:entry colname="col4">0.049</oasis:entry>
         <oasis:entry colname="col5">0.029</oasis:entry>
         <oasis:entry colname="col6">0.046</oasis:entry>
         <oasis:entry colname="col7">0.035</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">GW orog 1.0 %</oasis:entry>
         <oasis:entry colname="col2">0.054</oasis:entry>
         <oasis:entry colname="col3">0.032</oasis:entry>
         <oasis:entry colname="col4">0.036</oasis:entry>
         <oasis:entry colname="col5">0.025</oasis:entry>
         <oasis:entry colname="col6">0.034</oasis:entry>
         <oasis:entry colname="col7">0.026</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">GW orog 10.0 %</oasis:entry>
         <oasis:entry colname="col2">0.044</oasis:entry>
         <oasis:entry colname="col3">0.022</oasis:entry>
         <oasis:entry colname="col4">0.048</oasis:entry>
         <oasis:entry colname="col5">0.033</oasis:entry>
         <oasis:entry colname="col6">0.051</oasis:entry>
         <oasis:entry colname="col7">0.040</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">GW orog 20.0 %</oasis:entry>
         <oasis:entry colname="col2">0.061</oasis:entry>
         <oasis:entry colname="col3">0.039</oasis:entry>
         <oasis:entry colname="col4">0.051</oasis:entry>
         <oasis:entry colname="col5">0.033</oasis:entry>
         <oasis:entry colname="col6">0.055</oasis:entry>
         <oasis:entry colname="col7">0.041</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">GW orog 30.0 %</oasis:entry>
         <oasis:entry colname="col2">0.048</oasis:entry>
         <oasis:entry colname="col3">0.032</oasis:entry>
         <oasis:entry colname="col4">0.046</oasis:entry>
         <oasis:entry colname="col5">0.031</oasis:entry>
         <oasis:entry colname="col6">0.054</oasis:entry>
         <oasis:entry colname="col7">0.036</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">GW orog 40.0 %</oasis:entry>
         <oasis:entry colname="col2">0.035</oasis:entry>
         <oasis:entry colname="col3">0.031</oasis:entry>
         <oasis:entry colname="col4">0.050</oasis:entry>
         <oasis:entry colname="col5">0.044</oasis:entry>
         <oasis:entry colname="col6">0.051</oasis:entry>
         <oasis:entry colname="col7">0.044</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">GW orog 50.0 %</oasis:entry>
         <oasis:entry colname="col2">0.053</oasis:entry>
         <oasis:entry colname="col3">0.028</oasis:entry>
         <oasis:entry colname="col4">0.049</oasis:entry>
         <oasis:entry colname="col5">0.040</oasis:entry>
         <oasis:entry colname="col6">0.046</oasis:entry>
         <oasis:entry colname="col7">0.043</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">clubb c1 1.0 %</oasis:entry>
         <oasis:entry colname="col2">0.041</oasis:entry>
         <oasis:entry colname="col3">0.049</oasis:entry>
         <oasis:entry colname="col4">0.040</oasis:entry>
         <oasis:entry colname="col5">0.029</oasis:entry>
         <oasis:entry colname="col6">0.043</oasis:entry>
         <oasis:entry colname="col7">0.029</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">clubb c1 3.0 %</oasis:entry>
         <oasis:entry colname="col2">0.053</oasis:entry>
         <oasis:entry colname="col3">0.034</oasis:entry>
         <oasis:entry colname="col4">0.058</oasis:entry>
         <oasis:entry colname="col5">0.045</oasis:entry>
         <oasis:entry colname="col6">0.056</oasis:entry>
         <oasis:entry colname="col7">0.050</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">clubb c1 5.0 %</oasis:entry>
         <oasis:entry colname="col2">0.049</oasis:entry>
         <oasis:entry colname="col3">0.031</oasis:entry>
         <oasis:entry colname="col4">0.050</oasis:entry>
         <oasis:entry colname="col5">0.036</oasis:entry>
         <oasis:entry colname="col6">0.054</oasis:entry>
         <oasis:entry colname="col7">0.037</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">clubb c1 10.0 %</oasis:entry>
         <oasis:entry colname="col2">0.044</oasis:entry>
         <oasis:entry colname="col3">0.029</oasis:entry>
         <oasis:entry colname="col4">0.046</oasis:entry>
         <oasis:entry colname="col5">0.032</oasis:entry>
         <oasis:entry colname="col6">0.048</oasis:entry>
         <oasis:entry colname="col7">0.036</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">zmconv c0 ocn 0.5 %</oasis:entry>
         <oasis:entry colname="col2">0.041</oasis:entry>
         <oasis:entry colname="col3">0.029</oasis:entry>
         <oasis:entry colname="col4">0.049</oasis:entry>
         <oasis:entry colname="col5">0.042</oasis:entry>
         <oasis:entry colname="col6">0.050</oasis:entry>
         <oasis:entry colname="col7">0.047</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">zmconv c0 ocn 1.0 %</oasis:entry>
         <oasis:entry colname="col2">0.045</oasis:entry>
         <oasis:entry colname="col3">0.046</oasis:entry>
         <oasis:entry colname="col4">0.063</oasis:entry>
         <oasis:entry colname="col5">0.038</oasis:entry>
         <oasis:entry colname="col6">0.051</oasis:entry>
         <oasis:entry colname="col7">0.050</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">zmconv c0 ocn 3.0 %</oasis:entry>
         <oasis:entry colname="col2">0.032</oasis:entry>
         <oasis:entry colname="col3">0.028</oasis:entry>
         <oasis:entry colname="col4">0.045</oasis:entry>
         <oasis:entry colname="col5">0.030</oasis:entry>
         <oasis:entry colname="col6">0.047</oasis:entry>
         <oasis:entry colname="col7">0.040</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">zmconv c0 ocn 5.0 %</oasis:entry>
         <oasis:entry colname="col2">0.040</oasis:entry>
         <oasis:entry colname="col3">0.040</oasis:entry>
         <oasis:entry colname="col4">0.041</oasis:entry>
         <oasis:entry colname="col5">0.034</oasis:entry>
         <oasis:entry colname="col6">0.046</oasis:entry>
         <oasis:entry colname="col7">0.039</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">opt-O1</oasis:entry>
         <oasis:entry colname="col2">0.060</oasis:entry>
         <oasis:entry colname="col3">0.032</oasis:entry>
         <oasis:entry colname="col4">0.053</oasis:entry>
         <oasis:entry colname="col5">0.036</oasis:entry>
         <oasis:entry colname="col6">0.058</oasis:entry>
         <oasis:entry colname="col7">0.040</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">fastmath</oasis:entry>
         <oasis:entry colname="col2">0.059</oasis:entry>
         <oasis:entry colname="col3">0.034</oasis:entry>
         <oasis:entry colname="col4">0.043</oasis:entry>
         <oasis:entry colname="col5">0.030</oasis:entry>
         <oasis:entry colname="col6">0.043</oasis:entry>
         <oasis:entry colname="col7">0.040</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Mean</oasis:entry>
         <oasis:entry colname="col2">0.048</oasis:entry>
         <oasis:entry colname="col3">0.033</oasis:entry>
         <oasis:entry colname="col4">0.048</oasis:entry>
         <oasis:entry colname="col5">0.035</oasis:entry>
         <oasis:entry colname="col6">0.049</oasis:entry>
         <oasis:entry colname="col7">0.040</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">95th %tile</oasis:entry>
         <oasis:entry colname="col2">0.060</oasis:entry>
         <oasis:entry colname="col3">0.047</oasis:entry>
         <oasis:entry colname="col4">0.059</oasis:entry>
         <oasis:entry colname="col5">0.044</oasis:entry>
         <oasis:entry colname="col6">0.056</oasis:entry>
         <oasis:entry colname="col7">0.050</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>False Negative Rates and Statistical Power: MVK and BH-FDR</title>
      <p id="d2e2233">To evaluate the magnitude of change that the tests can detect confidently, we again rely on bootstrapping following previous work <xref ref-type="bibr" rid="bib1.bibx23 bib1.bibx22 bib1.bibx20" id="paren.49"/>. For each tuning parameter change, 30 ensemble members each are drawn from the control and that tuning parameter change ensemble. The test statistic, <inline-formula><mml:math id="M89" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, is then computed for each uncorrected statistical test (K-S, M-W, and C-VM) and the BH-FDR corrected version of each, then the six (three uncorrected, three BH-FDR corrected) tests are conducted on the ensemble pair. This procedure is then repeated 1000 times. To illustrate the impact of progressively increasing the tuning parameter on the model climate, Fig. <xref ref-type="fig" rid="F3"/> shows the 95th percentile of <inline-formula><mml:math id="M90" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> from the 1000 hypothesis tests for each tuning parameter change for <monospace>effgw_oro</monospace>, <monospace>clubb_c1</monospace>, and <monospace>zmconv_c0_ocn</monospace>. As the magnitude of the tuning parameter change to the model increases, <inline-formula><mml:math id="M91" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> increases for both uncorrected and BH-FDR corrected approaches. To reiterate, an increase in <inline-formula><mml:math id="M92" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> indicates an increase in the number of fields rejecting the local null hypothesis, <inline-formula><mml:math id="M93" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>. Figure <xref ref-type="fig" rid="F3"/> also re-iterates the lower sensitivity of E3SM to the orographic gravity wave drag parameter than both the C1 parameter from the CLUBB cloud parameterization scheme, and C0 over ocean in the ZM convection scheme. Smaller percentage changes in <monospace>clubb_c1</monospace> and <monospace>zmconv_c0_ocn</monospace> result in large changes to <inline-formula><mml:math id="M94" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> as compared to <monospace>effgw_oro</monospace>, where larger percentage changes are needed for similar changes in <inline-formula><mml:math id="M95" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>. The K-S test, as expected from the control threshold estimation, rejects the fewest number of fields at each level, while the C-VM and M-W tests reject nearly the same number. However, the number of rejected fields above each test's respective control threshold appears similar.</p>
      <p id="d2e2316">Figure <xref ref-type="fig" rid="F3"/> also points towards the detectability of modifications to the model by the tests. For the <monospace>effgw_oro</monospace> and <monospace>clubb_c1</monospace> parameters, a 1 % change in their value does not result in a change in the simulated climate that is easily detectable by the tests since <inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">95</mml:mn></mml:mrow></mml:math></inline-formula> % of the bootstrapped ensemble pairs, <inline-formula><mml:math id="M97" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> is less than the critical value threshold of the tests (11 for K-S, 16 for M-W and C-VM, and one for BH-FDR). This indicates that 95 % of the 1000 bootstrap comparisons between the control and, for example, the 10 % change to <monospace>effgw_oro</monospace>, have 13 or fewer of the 117 output fields rejected for the K-S test, 19 or fewer fields for the C-VM test, and 20 or fewer for the M-W test, and 1, 2, and 2 of 117 output fields rejected for BH-FDR corrected versions for each. The 95 % level is chosen here as it is the inverse of our significance level <inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula>, and as such one could expect to see 5 % of bootstrap iterations failing the test on random chance, thus if fewer than 5 % have passed, this is a strong indicator that the two ensembles are significantly different.</p>
      <p id="d2e2360">Increasing <monospace>effgw_oro</monospace> to 10 %, <monospace>clubb_c1</monospace> to 3 %, and <monospace>zmconv_c0_ocn</monospace> to 1 % results in some of the bootstrap iterations having exceeded or met the critical value threshold for both MVK and BH-FDR. As the magnitude of tuning parameter change increases, the number of bootstrap iterations crossing the critical value threshold also increases.</p>

      <fig id="F3"><label>Figure 3</label><caption><p id="d2e2375">95th percentile of the number of fields with statistically significant differences from the control ensemble (reject the local null hypothesis, <inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> at the <inline-formula><mml:math id="M100" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula> significance level) on the <inline-formula><mml:math id="M101" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axis by percentage change in tuning parameter along the <inline-formula><mml:math id="M102" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis. Solid green, blue, and orange lines represent K-S, C-VM, and M-W tests respectively, dashed represents BH-FDR versions of each in the same color. The dashed horizontal lines represent global null hypothesis critical value thresholds for K-S, C-VM, and M-W tests (green, orange respectively) and BH-FDR (black).</p></caption>
          <graphic xlink:href="https://esd.copernicus.org/articles/17/23/2026/esd-17-23-2026-f03.png"/>

        </fig>

      <p id="d2e2421">A formal estimate of the statistical power (<inline-formula><mml:math id="M103" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula>, rate of correctly rejecting a false null hypothesis), also representative of the false negative rates (<inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi>P</mml:mi></mml:mrow></mml:math></inline-formula>, incorrectly accepting a false null hypothesis), of these tests is illustrated in Fig. <xref ref-type="fig" rid="F4"/>. It shows the number of bootstrap iterations where the tests correctly reject <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> at a significance level of <inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula> and is indicative of the likelihood of the tests detecting a modification to the model. <inline-formula><mml:math id="M107" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> can be computed by dividing the ordinate (<inline-formula><mml:math id="M108" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> axis) values by the total number of bootstrap iterations (1000). Similar to Fig. <xref ref-type="fig" rid="F3"/>, Fig. <xref ref-type="fig" rid="F4"/> shows that as the magnitude of a tuning parameter is increased, the number of bootstrap iterations rejecting <inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> also increases. For a 1 % change to <monospace>effgw_oro</monospace> or <monospace>clubb_c1</monospace> only about 30–40 bootstrap iterations reject <inline-formula><mml:math id="M110" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, which implies that there is only a 30–40 out of a 1000 (3 %–4 %) chance that a change of this magnitude could be detected by the tests. At a 0.5 % change to <monospace>zmconv_c0_ocn</monospace>, about 60–70 bootstrap iterations reject <inline-formula><mml:math id="M111" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, implying a 6 %–7 % chance of detecting a change of this magnitude using these tests. At the <inline-formula><mml:math id="M112" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula> significance level it is expected that <inline-formula><mml:math id="M113" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">5</mml:mn><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="italic">%</mml:mi></mml:mrow></mml:math></inline-formula> of the tests will be rejected by random chance. Increasing <monospace>effgw_oro</monospace> to 10 % results in more than 50 bootstrap iterations (5 %) rejecting <inline-formula><mml:math id="M114" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, indicating that it unlikely to be caused by random chance at the 0.05 significance level, it still exhibits a low likelihood of being detected by the tests (about 60 and 80 out of a 1000 chance for MVK and BH-FDR tests respectively). As the magnitude of change to <monospace>effgw_oro</monospace> is increased, the likelihood of detecting a change by the tests increases. For a change of 40 % to <monospace>effgw_oro</monospace> there is a <inline-formula><mml:math id="M115" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">90</mml:mn></mml:mrow></mml:math></inline-formula> % chance of being detected by both the tests and it reaches nearly a 100 % for a larger change. Similarly, for <monospace>clubb_c1</monospace>, a change of 5 % results in greater than 80 % (90 %) chance of being detected by the MVK (BH-FDR) test, and nearly a 100 % chance of detection for a change of about 10 % change to its magnitude.</p>

      <fig id="F4"><label>Figure 4</label><caption><p id="d2e2591">Number of bootstrap iterations that reject the global null hypothesis (<inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) out of 1000 bootstrap iterations for changes in tuning parameter <monospace>effgw_oro</monospace> (top), <monospace>clubb_c1</monospace> (middle), and <monospace>zmconv_c0_ocn</monospace> (bottom). Dashed black lines represent 5 %–95 % of all bootstraps.</p></caption>
          <graphic xlink:href="https://esd.copernicus.org/articles/17/23/2026/esd-17-23-2026-f04.png"/>

        </fig>

      <p id="d2e2620">Overall, BH-FDR approach exhibits greater statistical power than MVK for almost all tuning parameter changes, allowing increased confidence in detecting changes, adding to its advantages. This is consistent with previous works <xref ref-type="bibr" rid="bib1.bibx3 bib1.bibx41" id="paren.50"/>, that suggest that BH-FDR approach generally exhibits greater power than bootstrapping methods. Figure <xref ref-type="fig" rid="F5"/> shows the power difference between uncorrected and BH-FDR corrected approaches. For the <monospace>clubb_c1</monospace> and <monospace>zmconv_c0_ocn</monospace> parameters, there is an increase in power for all tests at all parameter changes, apart from the very smallest and largest percent changes (at the largest no power increase is possible as the uncorrected approach already rejects all iterations). Both the C-VM and M-W tests have power increases for the <monospace>effwg_oro</monospace> parameter changes, but the K-S test has near 0 or power losses except at largest parameter changes.</p>
      <p id="d2e2637">For the smallest parameter change to <monospace>effgw_oro</monospace>, where MVK test exhibits greater power than the BH-FDR approach, it is possible that at the 1 % change to <monospace>effgw_oro</monospace>, the simulated climate is not different in a meaningful way. This means that the spread of simulated climates (measured by global averages of the output fields) generated by adding a random perturbation to the temperature field does not differ significantly from the spread of the simulated climates with a 1 % change to <monospace>effgw_oro</monospace>. Sampling errors may also be playing a role, given the small magnitude of change. Increasing the ensemble sizes can significantly increase the power of the tests as shown for MVK <xref ref-type="bibr" rid="bib1.bibx23" id="paren.51"/>, but operationally add to the computational cost. This trade-off between false negative rates and ensemble sizes is a decision left to model developers and code integrators. The power analysis here provides them with some reference to interpret the test results. For instance, if a non-bit-for-bit change passes the test (with the ensemble size of say, 30), developers can infer that its impact is likely smaller than a 5 % change in C1 parameter of CLUBB, which can be detected at a high confidence by the tests. This contextual comparison helps determine whether to accept or investigate a change further and guides the selection of ensemble size needed to detect changes of interest. In the future, we will expand our power analysis to include other tuning parameter changes to better inform developers, integrators and domain scientists using the tests.</p>

      <fig id="F5"><label>Figure 5</label><caption><p id="d2e2655">Additional number of bootstrap iterations that reject the global null hypothesis (<inline-formula><mml:math id="M117" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) out of 1000 bootstrap iterations for changes in tuning parameter <monospace>effgw_oro</monospace> (top), <monospace>clubb_c1</monospace> (middle), and <monospace>zmconv_c0_ocn</monospace> (bottom) when BH-FDR is applied for each statistical test, colors as in Fig. <xref ref-type="fig" rid="F4"/>.</p></caption>
          <graphic xlink:href="https://esd.copernicus.org/articles/17/23/2026/esd-17-23-2026-f05.png"/>

        </fig>

</sec>
<sec id="Ch1.S4.SS4">
  <label>4.4</label><title>Optimization changes</title>
      <p id="d2e2694">Two simulation ensembles (“opt-O1” and “fastmath”) were performed to apply the uncorrected and BH-FDR approaches to evaluate sensitivity of model results to compiler optimization flags, which are involved in the optimization of the underlying mathematics rather than model tuning parameters designed to account for varying physical processes. Bootstrapping procedures, similar to those described in the last section indicate that these optimizations do not have a significant impact on solution reproducibility (Fig. <xref ref-type="fig" rid="F6"/>). Between 10–31 out of the 1000 bootstrap iterations reject <inline-formula><mml:math id="M118" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> using the uncorrected statistical tests for both “opt-O1” and “fastmath” at the significance level of 0.05. After application of the BH-FDR approach, even fewer bootstrap iterations reject <inline-formula><mml:math id="M119" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, exhibiting a 0.5 %–0.7 % chance to detect a change to “O1” and a 2.3 %–2.9 % chance to detect a change in the “fastmath” flag. Reducing optimizations has been shown to yield climate reproducibility in previous studies as well <xref ref-type="bibr" rid="bib1.bibx2 bib1.bibx21" id="paren.52"/>, similar to our result that the simulated climate of “opt-O1” is statistically indistinguishable from that of the control ensemble that uses “-O3” optimizations.</p>

      <fig id="F6" specific-use="star"><label>Figure 6</label><caption><p id="d2e2726">Number of bootstrap iterations that reject the global null hypothesis (<inline-formula><mml:math id="M120" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) for opt-O1 (left), and fastmath (right) ensembles, when compared to the control ensemble, out of 1000 bootstrap iterations for each statistical test, uncorrected in blue and B-H FDR (orange).</p></caption>
          <graphic xlink:href="https://esd.copernicus.org/articles/17/23/2026/esd-17-23-2026-f06.png"/>

        </fig>

      <p id="d2e2746">The aggressive optimizations enabled by “-fp-model=fast” are expected to decrease solution accuracy <xref ref-type="bibr" rid="bib1.bibx27 bib1.bibx7" id="paren.53"/>, however, in this set of ensembles they do not significantly alter the simulated climate. As previously mentioned, the “-fp-model=fast” option is already a default for several source files, thus this simulation ensemble tests its use only for those additional source files in EAM. This is different from testing the impact of using it across the model as a whole as in <xref ref-type="bibr" rid="bib1.bibx21" id="text.54"/>. Acceptance of <inline-formula><mml:math id="M121" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> by the tests indicate that including the option for all of EAM, instead of selectively as is the default, does not result in statistically distinguishable solutions. This indicates that the “-fp-model=fast” optimization could be applied more generally throughout EAM code base. However, for higher resolutions or fully-coupled ensembles, this result may not apply as these configurations may respond to perturbations differently than ultra-low resolution model used here. Thus, further examination of its applicability may be required. </p>
</sec>
<sec id="Ch1.S4.SS5">
  <label>4.5</label><title>Standard resolution test</title>
      <p id="d2e2776">A test was also conducted using the USAF HPC11 using the same model version (E3SM v2.1), but at a higher resolution, called “standard-resolution” (called ne30pg2, <inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">1.0</mml:mn></mml:mrow></mml:math></inline-formula>° atmosphere) Two ensembles of 30 members each were constructed, a control ensemble with no parameter modification from default, and an ensemble where the <monospace>clubb_c1</monospace> parameter had been increased by 5 % from 2.4 to 2.52. Table <xref ref-type="table" rid="T3"/> details the number of output fields which are rejected at <inline-formula><mml:math id="M123" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula> level for both uncorrected and BH-FDR corrected methods for each statistical test. Though a large ensemble was not conducted to determine an empirical global rejection threshold for the uncorrected tests, it is reasonable to assume that following the BH-FDR correction, that any field rejection results in a global null hypothesis rejection as in the “ultra-low” resolution ensembles. This means that for all tests, the two simulated climates are determined to be significantly different. Figure <xref ref-type="fig" rid="F7"/> shows that the map of grid box averaged cloud liquid amount is visually different on the order of <inline-formula><mml:math id="M124" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">6</mml:mn></mml:mrow></mml:msup><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mrow class="unit"><mml:mi mathvariant="normal">kg</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msup><mml:mi mathvariant="normal">kg</mml:mi><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:math></inline-formula>, only one order of magnitude smaller than the field itself in the ensemble mean depending on the location. The largest differences appear mainly over the Northern Hemisphere, though smaller differences are noticeable in the Southern Ocean. Grid box-wise statistical tests are not conducted for this framework, but the field appears visually different in the mean sense, and is statistically distinct as confirmed by the statistical tests used.</p>

<table-wrap id="T3"><label>Table 3</label><caption><p id="d2e2845">Number of rejected fields in the “standard resolution” ensemble comparison for each statistical test, for uncorrected and BH-FDR corrected methods.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Statistical</oasis:entry>
         <oasis:entry colname="col2">Uncorrected</oasis:entry>
         <oasis:entry colname="col3">BH-FDR corrected</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">test</oasis:entry>
         <oasis:entry colname="col2">rejections</oasis:entry>
         <oasis:entry colname="col3">rejections</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">K-S</oasis:entry>
         <oasis:entry colname="col2">29</oasis:entry>
         <oasis:entry colname="col3">22</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">M-W</oasis:entry>
         <oasis:entry colname="col2">32</oasis:entry>
         <oasis:entry colname="col3">27</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">C-VM</oasis:entry>
         <oasis:entry colname="col2">35</oasis:entry>
         <oasis:entry colname="col3">25</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <fig id="F7"><label>Figure 7</label><caption><p id="d2e2927">Map of ensemble mean cloud liquid amount for the ne30 resolution control ensemble (top) and the difference to the perturbed ensemble mean (bottom) in kg kg<sup>−1</sup>.</p></caption>
          <graphic xlink:href="https://esd.copernicus.org/articles/17/23/2026/esd-17-23-2026-f07.png"/>

        </fig>

      <p id="d2e2949">The annual global means for each ensemble member are plotted against one another in Fig. <xref ref-type="fig" rid="F8"/>. This visualizes differences in the distributions of the cloud liquid amount field which are used by the statistical tests to generate <inline-formula><mml:math id="M126" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> values. For a field which has no statistically significant differences in distributions between experiment (<monospace>clubb_c1</monospace> + 5.0 %) and Control, the quantiles would fall along the <inline-formula><mml:math id="M127" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>:</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> line in the <inline-formula><mml:math id="M128" display="inline"><mml:mi>Q</mml:mi></mml:math></inline-formula>-<inline-formula><mml:math id="M129" display="inline"><mml:mi>Q</mml:mi></mml:math></inline-formula> plot, as would the probabilities along the same line in the <inline-formula><mml:math id="M130" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula>-<inline-formula><mml:math id="M131" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> plot, the histogram bars would be of similar heights, and the cumulative distribution functions would have little distance between them. This is not the case for the cloud liquid amount, and its local null hypothesis is rejected based on all tests and methods.</p>

      <fig id="F8"><label>Figure 8</label><caption><p id="d2e3007">Cumulative distribution function for the cloud liquid amount field, as in Fig. <xref ref-type="fig" rid="F7"/>. Both control and test ensemble annual global means are normalized together.</p></caption>
          <graphic xlink:href="https://esd.copernicus.org/articles/17/23/2026/esd-17-23-2026-f08.png"/>

        </fig>

</sec>
<sec id="Ch1.S4.SS6">
  <label>4.6</label><title>Ensemble size</title>
      <p id="d2e3026">The size of the sub-ensemble selected of 30 was chosen based on the previous results of <xref ref-type="bibr" rid="bib1.bibx24 bib1.bibx23" id="text.55"/> as a balance between statistical power and computational efficiency. Figure <xref ref-type="fig" rid="F9"/> shows an increasing power for larger sub-ensemble size selections, with the BH-FDR corrected approaches exhibiting larger power than the uncorrected counterparts. Here, the choice of 30 ensemble member comparisons appears an appropriate balance between statistical power and computational time.</p>

      <fig id="F9"><label>Figure 9</label><caption><p id="d2e3036">Power of statistical tests as in Fig. <xref ref-type="fig" rid="F4"/> on the <monospace>clubb_c1</monospace> 5 % change to control ensemble comparison, but for varying sub-ensemble size selection (<inline-formula><mml:math id="M132" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> axis). Solid curves indicate uncorrected methods, dashed lines of the same color indicate BH-FDR corrected approach for each test.</p></caption>
          <graphic xlink:href="https://esd.copernicus.org/articles/17/23/2026/esd-17-23-2026-f09.png"/>

        </fig>

      <p id="d2e3057">In the other figures of this study, the results are presented with an ensemble size of 30, to match the operational constraints so results are directly applicable on the nightly testing.</p>
</sec>
<sec id="Ch1.S4.SS7">
  <label>4.7</label><title>Operational results</title>
      <p id="d2e3068">Figure <xref ref-type="fig" rid="F10"/> shows the time series of <inline-formula><mml:math id="M133" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, or the number of variables rejecting <inline-formula><mml:math id="M134" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, for MVK and BH-FDR, for a period of a few weeks after BH-FDR was implemented and included in nightly testing in late September last year. The additional statistical tests were not performed on these ensembles, and are thus not able to be included here, as the nightly testing output is not archived. The model maintained bit-for-bit reproducibility during these weeks. Bit-for-bit reproducibility is ascertained by a suite of bit-for-bit tests that are run each night, testing the model under a variety of conditions. When these pass, the model is bit-for-bit with previous results, and a test fail by MVK and BH-FDR on these days is thus known to be a false positive. Figure <xref ref-type="fig" rid="F10"/> shows that operationally BH-FDR has a reduced false positive rate of 1.9 % (1 of 53 tests) as compared to 7.5 % (4 of 53 tests) for MVK. <inline-formula><mml:math id="M135" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> varies between 0 to a maximum of 15 (on two individual days) for MVK, while <inline-formula><mml:math id="M136" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> is either 0 or 1 (on one occasion) for BH-FDR. This solitary global null hypothesis rejection using BH-FDR does not occur at the same time as a global rejection using uncorrected <inline-formula><mml:math id="M137" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> values, indicating that a systematic change in its statistics may not have occurred. In which case both tests may be expected to fail. The BH-FDR approach yields a failing overall result when one particular <inline-formula><mml:math id="M138" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> value is very small. In this case the <monospace>soa_a1_SRF</monospace>, a secondary organic aerosol field, was rejected very strongly, which meant the BH-FDR method rejected the global null hypothesis. The rate of global rejections using MVK (7.5 %) is higher than the targeted 5 %, which is reduced to under 2 % using the FDR corrected <inline-formula><mml:math id="M139" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> value threshold, as is expected as the correction controls for the false discovery rate when using BH-FDR. The higher false positive rate of MVK can be associated with the small sample size (53 d of testing). In the future, we plan to expand the ensemble size of the control ensembles to derive the null distribution and evaluate its impact on operational false positive rates of MVK.</p>

      <fig id="F10"><label>Figure 10</label><caption><p id="d2e3134">Number of rejected variables (of 120) for nightly tests of E3SM. The teal line shows rejection based on MVK <inline-formula><mml:math id="M140" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> values, orange shows number of rejections based on B-H FDR <inline-formula><mml:math id="M141" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> value thresholds.</p></caption>
          <graphic xlink:href="https://esd.copernicus.org/articles/17/23/2026/esd-17-23-2026-f10.png"/>

        </fig>

      <p id="d2e3157">Operationally, the MVK test has recently been successful in identifying two bugs committed to the model's code, which were thought by developers to be non bit-for-bit, but not climate changing. The test returned a “fail” result following both changes, and the changes were reverted and eventually re-worked into non climate changing code alterations. This unintentional climate changing update was also captured by the BH-FDR approach and indicates its ability to capture erroneous alterations in addition to its ability to capture tuning parameter changes. In the first case, developers introduced changes related to aqueous-phase chemistry, specifically involving reactions tied to cloud water and trace gases. These updates were considered minor adjustments to internal model behavior and were not expected to alter the simulated climate. However, the tests flagged the update as a baseline failure, revealing statistically significant differences across many atmospheric variables. Further investigation showed that the change affected cloud-aerosol interactions and subsequently altered radiative balance, demonstrating how a seemingly minor, but non-bit-for-bit change in chemical processes led to broader simulated climate impacts. In the second case, ozone chemistry configuration changes were merged, aiming to improve numerical consistency in offline and online chemistry calculations and were also expected to be non-climate changing. However, MVK and BH-FDR detected a statistically significant difference, which was traced to elevated tropospheric ozone concentrations resulting from the changes. The alteration impacted radiative forcing enough to shift the ensemble mean state, and was also subsequently corrected. These examples illustrate that the BH-FDR method has statistical power in the event of erroneous code changes in addition to its power in detecting perturbed parameters as does the MVK, thus its usefulness is not degraded.</p>
</sec>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <label>5</label><title>Summary and discussion</title>
      <p id="d2e3170">This study presents a new approach, BH-FDR, to evaluate statistical solution reproducibility of EAM after unintended non-bit-for-bit changes are introduced. BH-FDR improves on the existing MVK test and applies a false discovery rate correction to control Type I error inflation in multi-testing scenarios. While the original MVK approach relies on computationally expensive bootstrapping to determine critical value thresholds, BH-FDR offers a theoretically grounded and operationally simpler alternative. This computationally expensive threshold finding, where the ensembles represent <inline-formula><mml:math id="M142" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">3700</mml:mn></mml:mrow></mml:math></inline-formula> node hours of computation time for ensemble generation, is demonstrated to be effectively eliminated, as the BH-FDR approach for all statistical tests can use a threshold of 1, while maintaining nearly the same, or improving statistical power.</p>
      <p id="d2e3183">Our evaluation using a comprehensive suite of ensembles, including both parameter perturbations and compiler optimization changes, demonstrates that BH-FDR approach maintains or improves the statistical power of the MVK test, while significantly reducing false positive rates. Notably, the BH-FDR approach eliminates the need for re-calibrating critical value thresholds after major model revisions. Operational implementation of this method in nightly E3SM testing has further validated its utility, showing a reduction in false positives from 7.5 % to 1.9 %. Overall, the BH-FDR approach enhances the robustness, accuracy, and efficiency of statistical testing for climate model reproducibility.</p>
      <p id="d2e3186">Additional tests were performed with the Cramér-von Mises and Mann-Whitney <inline-formula><mml:math id="M143" display="inline"><mml:mi>U</mml:mi></mml:math></inline-formula> tests to add a broader perspective on the applicability of the BH-FDR method. For uncorrected methods, the K-S test had the highest power and equivalent false positive rates to the C-VM and M-W tests, but in using the BH-FDR method, the power of those tests increased beyond that of the K-S test. Primarily, this indicates that the BH-FDR method works well for multiple different statistical tests and is valid so long as the underlying statistical test is valid for the data being examined. We plan to add these additional statistical tests to the nightly suite and continue to examine ways to usefully combine the results of these different tests to evaluate if two simulation ensembles are statistically similar.</p>
      <p id="d2e3196">We also examined the sensitivity of the results to ensemble size selection. Part of the procedure applied here to find the power of the various tests involves comparing two sub-ensembles each drawn from two larger ensembles at random. The power of all three statistical tests both for uncorrected and BH-FDR corrected methods scales with sub-ensemble size. There is, however, a trade-off between computational efficiency and statistical power. As an example, in this configuration with “ultra-low” resolution on the Chrysalis machine, one ensemble member runs on a single node in approximately 15 wall-seconds per model day or about 1 h 45 min per 14-month simulation. Each additional ensemble member then adds a node to the computational requirements, so a 30 member ensemble uses 30 nodes for 1 h 45 min, or approximately 53 node-hours. We believe that an ensemble size of 30 strikes the right balance between statistical power and computational costs for routine nightly testing, but a developer is able to choose to increase this size to enhance the statistical power as needed.</p>
      <p id="d2e3200">A caveat of our study here is that BH-FDR (and MVK as well) has been rigorously applied to and evaluated for the ultra-low resolution version of the model, which is not used in practical applications, and only a single test was performed at a higher resolution. The consistency of test results across two different model resolutions provides evidence that the test results with the ultra-low resolution model hold for standard and high resolution model configurations used for production runs. Higher resolution models resolve finer scale processes, which can effect numerical sensitivity, internal variability and process feedbacks of the model. In the future, we plan to identify the underlying reasons and scenarios in which test results may or may not remain consistent across resolutions. Nonetheless, an earlier unpublished work found that the results of MVK applied to ultra-low resolution ensembles and MVK applied to standard resolution ensembles were consistent when evaluating a port of an earlier version of E3SM to a new machine at the National Energy Research Scientific Computing Center (NERSC). Further, we will explore enhancing computational feasibility of the tests by using shorter run times of simulation ensembles <xref ref-type="bibr" rid="bib1.bibx28" id="paren.56"/>, allowing for routine testing with higher resolution models.</p>
      <p id="d2e3206">In addition to its application in traditional Earth system models, the BH-FDR-based statistical testing framework has potential for assessing the reproducibility of AI-based climate models, which are rapidly gaining prominence in climate science. Recent developments such as ClimateBench <xref ref-type="bibr" rid="bib1.bibx39" id="paren.57"/>, FourCastNet <xref ref-type="bibr" rid="bib1.bibx29" id="paren.58"/>, and Pangu-Weather <xref ref-type="bibr" rid="bib1.bibx5" id="paren.59"/> demonstrate the capabilities of deep learning models to emulate or replace components of physics-based models with significant computational advantages allowing the generation of very large ensembles at very low computational costs. However, verifying the reliability and reproducibility of these models presents unique challenges. AI models often involve stochastic elements in training, sensitivity to floating-point precision, and reliance on hardware-specific optimizations, all of which can lead to variability in output across runs. Standard bit-for-bit reproducibility tests are inadequate in this context, and statistical frameworks like MVK or BH-FDR could serve as robust alternatives to assess whether differences in AI model outputs are statistically meaningful or within expected variability. Prior work has highlighted the need for principled evaluation methods tailored to the probabilistic nature of machine learning in scientific applications <xref ref-type="bibr" rid="bib1.bibx31 bib1.bibx9" id="paren.60"/>, and integrating ensemble-based hypothesis testing into AI model workflows could be a step toward more rigorous, interpretable, and trustworthy deployment of AI systems in operational climate modeling. Ultra-low resolution models, like the one used here, can also be used to create very large ensembles at low computational cost allowing comparisons with their AI surrogate large ensembles. We plan to conduct such evaluations in the near future.</p>
</sec>

      
      </body>
    <back><notes notes-type="codeavailability"><title>Code availability</title>

      <p id="d2e3225">The code for this work can be found at <uri>https://github.com/mkstratos/detectable_climate</uri> (last access: 24 October 2025) <xref ref-type="bibr" rid="bib1.bibx18" id="paren.61"><named-content content-type="pre"><ext-link xlink:href="https://doi.org/10.5281/zenodo.17438094" ext-link-type="DOI">10.5281/zenodo.17438094</ext-link>,</named-content></xref></p>
  </notes><notes notes-type="dataavailability"><title>Data availability</title>

      <p id="d2e3240">The bootstrap data from each comparison is available at <ext-link xlink:href="https://doi.org/10.5281/zenodo.17438071" ext-link-type="DOI">10.5281/zenodo.17438071</ext-link> <xref ref-type="bibr" rid="bib1.bibx19" id="paren.62"/></p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d2e3251">MK and SM developed the methodology. MK wrote the code, conducted the simulations, and wrote the first draft of the manuscript. SM supervised the project and contributed writing to the final manuscript.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d2e3257">The contact author has declared that neither of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d2e3263">Any subjective views or opinions that might be expressed in the paper do not necessarily represent the views of the U.S. Department of Energy or the United States Government. Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.</p>
  </notes><notes notes-type="sistatement"><title>Special issue statement</title>

      <p id="d2e3272">This article is part of the special issue “Theoretical and computational aspects of ensemble design, implementation, and interpretation in climate science (ESD/GMD/NPG inter-journal SI)”. It is not associated with a conference.</p>
  </notes><ack><title>Acknowledgements</title><p id="d2e3278">This research was supported as part of the Energy Exascale Earth System Model (E3SM) project, funded by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research. The authors also gratefully acknowledge the computing resources provided on Blues, a high-performance computing cluster operated by the Laboratory Computing Resource Center at Argonne National Laboratory. This research used resources of the Oak Ridge Leadership Computing Facility, which is a DOE Office of Science User Facility supported under Contract DE-AC05-00OR22725. The authors also acknowledge the numerous open-source libraries on which this work depends, <xref ref-type="bibr" rid="bib1.bibx14" id="text.63"/>, <xref ref-type="bibr" rid="bib1.bibx36" id="text.64"/>, <xref ref-type="bibr" rid="bib1.bibx15" id="text.65"/>, <xref ref-type="bibr" rid="bib1.bibx8" id="text.66"/>, <xref ref-type="bibr" rid="bib1.bibx26" id="text.67"/>, <xref ref-type="bibr" rid="bib1.bibx16" id="text.68"/>, <xref ref-type="bibr" rid="bib1.bibx38" id="text.69"/>, <xref ref-type="bibr" rid="bib1.bibx34" id="text.70"/>. The authors also acknowledge the seven anonymous reviewers, their comments have made this a more robust investigation.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d2e3309">This research was supported as part of the Energy Exascale Earth System Model (E3SM) project, funded by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research.</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d2e3315">This paper was edited by Irina Tezaur and reviewed by Teo Price-Broncucia and six anonymous referees.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Anderson(1962)</label><mixed-citation>Anderson, T. W.: On the Distribution of the Two-Sample Cramer-von Mises Criterion, Ann. Math. Statist., 33, 1148–1159, <ext-link xlink:href="https://doi.org/10.1214/aoms/1177704477" ext-link-type="DOI">10.1214/aoms/1177704477</ext-link>, 1962.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Baker et al.(2015)</label><mixed-citation>Baker, A. H., Hammerling, D. M., Levy, M. N., Xu, H., Dennis, J. M., Eaton, B. E., Edwards, J., Hannay, C., Mickelson, S. A., Neale, R. B., Nychka, D., Shollenberger, J., Tribbia, J., Vertenstein, M., and Williamson, D.: A new ensemble-based consistency test for the Community Earth System Model (pyCECT v1.0), Geosci. Model Dev., 8, 2829–2840, <ext-link xlink:href="https://doi.org/10.5194/gmd-8-2829-2015" ext-link-type="DOI">10.5194/gmd-8-2829-2015</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Benjamini and Hochberg(1995)</label><mixed-citation>Benjamini, Y. and Hochberg, Y.: Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing, Journal of the Royal Statistical Society: Series B (Methodological), 57, 289–300, <ext-link xlink:href="https://doi.org/10.1111/j.2517-6161.1995.tb02031.x" ext-link-type="DOI">10.1111/j.2517-6161.1995.tb02031.x</ext-link>, 1995.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Benjamini and Yekutieli(2001)</label><mixed-citation>Benjamini, Y. and Yekutieli, D.: The control of the false discovery rate in multiple testing under dependency, The Annals of Statistics, 29, 1165–1188, <ext-link xlink:href="https://doi.org/10.1214/aos/1013699998" ext-link-type="DOI">10.1214/aos/1013699998</ext-link>, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Bi et al.(2022)</label><mixed-citation> Bi, K.,  Xie, L.,  Zhang, H., Chen, X., Gu, X., and Tian, Q.: Accurate medium-range global weather forecasting with 3D neural networks, Nature, 610, 87–93, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Burrell et al.(2020)</label><mixed-citation>Burrell, A. L., Evans, J. P., and De Kauwe, M. G.: Anthropogenic climate change has driven over 5 million km<sup>2</sup> of drylands towards desertification, Nature Communications, 11, 3853, <ext-link xlink:href="https://doi.org/10.1038/s41467-020-17710-7" ext-link-type="DOI">10.1038/s41467-020-17710-7</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Büttner et al.(2024)</label><mixed-citation>Büttner, M., Alt, C., Kenter, T., Köstler, H., Plessl, C., and Aizinger, V.: Enabling Performance Portability for Shallow Water Equations on CPUs, GPUs, and FPGAs with SYCL, in: Proceedings of the Platform for Advanced Scientific Computing Conference, PASC '24, Association for Computing Machinery, New York, NY, USA, ISBN 9798400706394, <ext-link xlink:href="https://doi.org/10.1145/3659914.3659925" ext-link-type="DOI">10.1145/3659914.3659925</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Dask Development Team(2016)</label><mixed-citation>Dask Development Team: Dask: Library for dynamic task scheduling, <uri>http://dask.pydata.org</uri> (last access: 24 October 2025), 2016.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Dueben and Bauer(2018)</label><mixed-citation>Dueben, P. D. and Bauer, P.: Challenges and design choices for global weather and climate models based on machine learning, Geosci. Model Dev., 11, 3999–4009, <ext-link xlink:href="https://doi.org/10.5194/gmd-11-3999-2018" ext-link-type="DOI">10.5194/gmd-11-3999-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>E3SM Project(2023)</label><mixed-citation>E3SM Project, D.: Energy Exascale Earth System Model v2.1.0, DOE Code [software],  <ext-link xlink:href="https://doi.org/10.11578/E3SM/dc.20230110.5" ext-link-type="DOI">10.11578/E3SM/dc.20230110.5</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Eidhammer et al.(2024)</label><mixed-citation>Eidhammer, T., Gettelman, A., Thayer-Calder, K., Watson-Parris, D., Elsaesser, G., Morrison, H., van Lier-Walqui, M., Song, C., and McCoy, D.: An extensible perturbed parameter ensemble for the Community Atmosphere Model version 6, Geosci. Model Dev., 17, 7835–7853, <ext-link xlink:href="https://doi.org/10.5194/gmd-17-7835-2024" ext-link-type="DOI">10.5194/gmd-17-7835-2024</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Gentle(2003)</label><mixed-citation>Gentle, J. E.: Random Number Generation and Monte Carlo Methods, Springer-Verlag, ISBN 0387001786, <ext-link xlink:href="https://doi.org/10.1007/b97336" ext-link-type="DOI">10.1007/b97336</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Golaz et al.(2022)</label><mixed-citation>Golaz, J.-C., Van Roekel, L. P., Zheng, X., Roberts, A. F., Wolfe, J. D., Lin, W., Bradley, A. M., Tang, Q., Maltrud, M. E., Forsyth, R. M., Zhang, C., Zhou, T., Zhang, K., Zender, C. S., Wu, M., Wang, H., Turner, A. K., Singh, B., Richter, J. H., Qin, Y., Petersen, M. R., Mametjanov, A., Ma, P.-L., Larson, V. E., Krishna, J., Keen, N. D., Jeffery, N., Hunke, E. C., Hannah, W. M., Guba, O., Griffin, B. M., Feng, Y., Engwirda, D., Di Vittorio, A. V., Dang, C., Conlon, L. M., Chen, C.-C.-J., Brunke, M. A., Bisht, G., Benedict, J. J., Asay-Davis, X. S., Zhang, Y., Zhang, M., Zeng, X., Xie, S., Wolfram, P. J., Vo, T., Veneziani, M., Tesfa, T. K., Sreepathi, S., Salinger, A. G., Reeves Eyre, J. E. J., Prather, M. J., Mahajan, S., Li, Q., Jones, P. W., Jacob, R. L., Huebler, G. W., Huang, X., Hillman, B. R., Harrop, B. E., Foucar, J. G., Fang, Y., Comeau, D. S., Caldwell, P. M., Bartoletti, T., Balaguru, K., Taylor, M. A., McCoy, R. B., Leung, L. R., and Bader, D. C.: The DOE E3SM Model Version 2: Overview of the Physical Model and Initial Model Evaluation, Journal of Advances in Modeling Earth Systems, 14, e2022MS003156, <ext-link xlink:href="https://doi.org/10.1029/2022MS003156" ext-link-type="DOI">10.1029/2022MS003156</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Harris et al.(2020)</label><mixed-citation>Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del Río, J. F., Wiebe, M., Peterson, P., Gérard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E.: Array programming with NumPy, Nature, 585, 357–362, <ext-link xlink:href="https://doi.org/10.1038/s41586-020-2649-2" ext-link-type="DOI">10.1038/s41586-020-2649-2</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Hoyer and Hamman(2017)</label><mixed-citation>Hoyer, S. and Hamman, J.: xarray: N-D labeled arrays and datasets in Python, Journal of Open Research Software, 5, <ext-link xlink:href="https://doi.org/10.5334/jors.148" ext-link-type="DOI">10.5334/jors.148</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Hunter(2007)</label><mixed-citation>Hunter, J. D.: Matplotlib: A 2D graphics environment, Computing in Science &amp; Engineering, 9, 90–95, <ext-link xlink:href="https://doi.org/10.1109/MCSE.2007.55" ext-link-type="DOI">10.1109/MCSE.2007.55</ext-link>, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Intel Corporation(2023)</label><mixed-citation>Intel Corporation: Intel Fortran Compiler Developer Guide and Reference, <uri>https://www.intel.com/content/www/us/en/docs/fortran-compiler/developer-guide-reference/2023-0/compiler-options-001.html</uri> (last access: 1 May 2025), 2023.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Kelleher and Mahajan(2025a)</label><mixed-citation>Kelleher, M. and Mahajan, S.: Detectable Climate (v1.1.0), Zenodo [code], <ext-link xlink:href="https://doi.org/10.5281/zenodo.17438094" ext-link-type="DOI">10.5281/zenodo.17438094</ext-link>, 2025a.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Kelleher and Mahajan(2025b)</label><mixed-citation>Kelleher, M. and Mahajan, S.: Detectable Climate Bootstrap Data (Version v2), Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.17438071" ext-link-type="DOI">10.5281/zenodo.17438071</ext-link>, 2025b.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Mahajan(2021)</label><mixed-citation>Mahajan, S.: Ensuring statistical reproducibility of ocean model simulations in the age of hybrid computing, in: Proceedings of the Platform for Advanced Scientific Computing Conference, PASC '21, Association for Computing Machinery, New York, NY, USA, ISBN 9781450385633, <ext-link xlink:href="https://doi.org/10.1145/3468267.3470572" ext-link-type="DOI">10.1145/3468267.3470572</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Mahajan et al.(2017)</label><mixed-citation>Mahajan, S., Gaddis, A. L., Evans, K. J., and Norman, M. R.: Exploring an Ensemble-Based Approach to Atmospheric Climate Modeling and Testing at Scale, international Conference on Computational Science, ICCS 2017, 12-14 June 2017, Zurich, Switzerland, Procedia Computer Science, 108, 735–744, <ext-link xlink:href="https://doi.org/10.1016/j.procs.2017.05.259" ext-link-type="DOI">10.1016/j.procs.2017.05.259</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Mahajan et al.(2019a)</label><mixed-citation> Mahajan, S., Evans, K. J., Kennedy, J. H., Xu, M., and Norman, M. R.: A multivariate approach to ensure statistical reproducibility of climate model simulations, in: Proceedings of the Platform for Advanced Scientific Computing Conference,  1–10, 2019a.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Mahajan et al.(2019b)</label><mixed-citation>Mahajan, S., Evans, K. J., Kennedy, J. H., Xu, M., Norman, M. R., and Branstetter, M. L.: Ongoing solution reproducibility of earth system models as they progress toward exascale computing, The International Journal of High Performance Computing Applications, 33, 784–790, <ext-link xlink:href="https://doi.org/10.1177/1094342019837341" ext-link-type="DOI">10.1177/1094342019837341</ext-link>, 2019b.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Mahajan et al.(2022)</label><mixed-citation> Mahajan, S., Tang, Q., Keen, N. D., Golaz, J.-C., and van Roekel, L. P.: Simulation of ENSO teleconnections to precipitation extremes over the United States in the high-resolution version of E3SM, Journal of Climate, 35, 3371–3393, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Mann and Whitney(1947)</label><mixed-citation>Mann, H. and Whitney, D. R.: On a Test of Whether one of Two Random Variables is Stochastically Larger than the Other., Ann. Math. Statist., 18, 50–60, <ext-link xlink:href="https://doi.org/10.1214/aoms/1177730491" ext-link-type="DOI">10.1214/aoms/1177730491</ext-link>, 1947.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>McKinney(2010)</label><mixed-citation>McKinney, W.: Data Structures for Statistical Computing in Python, in: Proceedings of the 9th Python in Science Conference, edited by: van der Walt, S. and Millman, J.,  56–61, <ext-link xlink:href="https://doi.org/10.25080/Majora-92bf1922-00a" ext-link-type="DOI">10.25080/Majora-92bf1922-00a</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Mielikainen et al.(2016)</label><mixed-citation>Mielikainen, J., Price, E., Huang, B., Huang, H.-L. A., and Lee, T.: GPU Compute Unified Device Architecture (CUDA)-based Parallelization of the RRTMG Shortwave Rapid Radiative Transfer Model, IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 9, 921–931, <ext-link xlink:href="https://doi.org/10.1109/JSTARS.2015.2427652" ext-link-type="DOI">10.1109/JSTARS.2015.2427652</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Milroy et al.(2018)</label><mixed-citation>Milroy, D. J., Baker, A. H., Hammerling, D. M., and Jessup, E. R.: Nine time steps: ultra-fast statistical consistency testing of the Community Earth System Model (pyCECT v3.0), Geosci. Model Dev., 11, 697–711, <ext-link xlink:href="https://doi.org/10.5194/gmd-11-697-2018" ext-link-type="DOI">10.5194/gmd-11-697-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>Pathak et al.(2022)</label><mixed-citation>Pathak, J., Subramanian, S., Harrington, P., Raja, S., Chattopadhyay, A.,  Mardani, M.,  Kurth, T.,  Hall, D.,  Li, Z.,  Azizzadenesheli, K.,  Hassanzadeh, P.,  Kashinath, K., and Anandkumar, A.: FourCastNet: A Global Data-driven High-resolution Weather Model using Adaptive Fourier Neural Operators, arXiv [preprint], <ext-link xlink:href="https://doi.org/10.48550/arXiv.2202.11214" ext-link-type="DOI">10.48550/arXiv.2202.11214</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>Qian et al.(2018)</label><mixed-citation>Qian, Y., Wan, H., Yang, B., Golaz, J.-C., Harrop, B., Hou, Z., Larson, V. E., Leung, L. R., Lin, G., Lin, W., Ma, P.-L., Ma, H.-Y., Rasch, P., Singh, B., Wang, H., Xie, S., and Zhang, K.: Parametric Sensitivity and Uncertainty Quantification in the Version 1 of E3SM Atmosphere Model Based on Short Perturbed Parameter Ensemble Simulations, Journal of Geophysical Research: Atmospheres, 123, 13046–13073, <ext-link xlink:href="https://doi.org/10.1029/2018JD028927" ext-link-type="DOI">10.1029/2018JD028927</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Rasp et al.(2020)</label><mixed-citation>Rasp, S., Dueben, P. D., Scher, S., Weyn, J. A., Mouatadid, S., and Thuerey, N.: WeatherBench: A benchmark dataset for data-driven weather forecasting, Journal of Advances in Modeling Earth Systems, 12, e2020MS002203, <ext-link xlink:href="https://doi.org/10.1029/2020MS002203" ext-link-type="DOI">10.1029/2020MS002203</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Renard et al.(2008)</label><mixed-citation>Renard, B., Lang, M., Bois, P., Dupeyrat, A., Mestre, O., Niel, H.,  Sauquet, E., Prudhomme, C., Parey, S., Paquet, E., Neppel, L., and Gailhard, J.: Regional methods for trend detection: Assessing field significance and regional consistency, Water Resources Research, 44, <ext-link xlink:href="https://doi.org/10.1029/2007WR006268" ext-link-type="DOI">10.1029/2007WR006268</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Rosinski and Williamson(1997)</label><mixed-citation>Rosinski, J. M. and Williamson, D. L.: The Accumulation of Rounding Errors and Port Validation for Global Atmospheric Models, SIAM Journal on Scientific Computing, 18, 552–564, <ext-link xlink:href="https://doi.org/10.1137/S1064827594275534" ext-link-type="DOI">10.1137/S1064827594275534</ext-link>, 1997.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>Seabold and Perktold(2010)</label><mixed-citation>Seabold, S. and Perktold, J.: statsmodels: Econometric and statistical modeling with Python, in: 9th Python in Science Conference, 28 June–3 July 2010, Austin, TX, USA, <ext-link xlink:href="https://doi.org/10.25080/Majora-92bf1922-012" ext-link-type="DOI">10.25080/Majora-92bf1922-012</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx35"><label>Ventura et al.(2004)</label><mixed-citation>Ventura, V., Paciorek, C. J., and Risbey, J. S.: Controlling the Proportion of Falsely Rejected Hypotheses when Conducting Multiple Tests with Climatological Data, Journal of Climate, 17, 4343–4356, <ext-link xlink:href="https://doi.org/10.1175/3199.1" ext-link-type="DOI">10.1175/3199.1</ext-link>, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>Virtanen et al.(2020)</label><mixed-citation>Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J., Polat, İ., Feng, Y., Moore, E. W., VanderPlas, J., Laxalde, D., Perktold, J., Cimrman, R., Henriksen, I., Quintero, E. A., Harris, C. R., Archibald, A. M., Ribeiro, A. H., Pedregosa, F., van Mulbregt, P., and SciPy 1.0 Contributors: SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python, Nature Methods, 17, 261–272, <ext-link xlink:href="https://doi.org/10.1038/s41592-019-0686-2" ext-link-type="DOI">10.1038/s41592-019-0686-2</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Wan et al.(2017)</label><mixed-citation>Wan, H., Zhang, K., Rasch, P. J., Singh, B., Chen, X., and Edwards, J.: A new and inexpensive non-bit-for-bit solution reproducibility test based on time step convergence (TSC1.0), Geosci. Model Dev., 10, 537–552, <ext-link xlink:href="https://doi.org/10.5194/gmd-10-537-2017" ext-link-type="DOI">10.5194/gmd-10-537-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Waskom(2021)</label><mixed-citation>Waskom, M. L.: seaborn: statistical data visualization, Journal of Open Source Software, 6, 3021, <ext-link xlink:href="https://doi.org/10.21105/joss.03021" ext-link-type="DOI">10.21105/joss.03021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Watson-Parris et al.(2022)</label><mixed-citation>Watson-Parris, D., Rao, Y., Olivié, D., Seland, Ø., Nowack, P., Camps-Valls, G., Stier, P., Bouabid, S., Dewey, M., Fons, E., Gonzalez, J., Harder, P., Jeggle, K., Lenhardt, J., Manshausen, P., Novitasari, M., Ricard, L., and Roesch, C.: ClimateBench v1.0: A Benchmark for Data-Driven Climate Projections, Journal of Advances in Modeling Earth Systems, 14,  <ext-link xlink:href="https://doi.org/10.1029/2021MS002954" ext-link-type="DOI">10.1029/2021MS002954</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx40"><label>Whan and Zwiers(2017)</label><mixed-citation> Whan, K. and Zwiers, F.: The impact of ENSO and the NAO on extreme winter precipitation in North America in observations and regional climate models, Climate Dynamics, 48, 1401–1411, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx41"><label>Wilks(2006)</label><mixed-citation>Wilks, D. S.: On “Field Significance” and the False Discovery Rate, Journal of Applied Meteorology and Climatology, 45, 1181–1189, <ext-link xlink:href="https://doi.org/10.1175/JAM2404.1" ext-link-type="DOI">10.1175/JAM2404.1</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx42"><label>Wilks(2016)</label><mixed-citation>Wilks, D. S.: “The Stippling Shows Statistically Significant Grid Points”: How Research Results are Routinely Overstated and Overinterpreted, and What to Do about It, Bulletin of the American Meteorological Society, 97, 2263–2273, <ext-link xlink:href="https://doi.org/10.1175/BAMS-D-15-00267.1" ext-link-type="DOI">10.1175/BAMS-D-15-00267.1</ext-link>, 2016. </mixed-citation></ref>
      <ref id="bib1.bibx43"><label>Zeman and Schär(2022)</label><mixed-citation>Zeman, C. and Schär, C.: An ensemble-based statistical methodology to detect differences in weather and climate model executables, Geosci. Model Dev., 15, 3183–3203, <ext-link xlink:href="https://doi.org/10.5194/gmd-15-3183-2022" ext-link-type="DOI">10.5194/gmd-15-3183-2022</ext-link>, 2022.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Enhanced climate reproducibility testing with false discovery rate correction</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Anderson(1962)</label><mixed-citation>
      
Anderson, T. W.: On the Distribution of the Two-Sample Cramer-von Mises
Criterion, Ann. Math. Statist., 33, 1148–1159,
<a href="https://doi.org/10.1214/aoms/1177704477" target="_blank">https://doi.org/10.1214/aoms/1177704477</a>, 1962.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Baker et al.(2015)</label><mixed-citation>
      
Baker, A. H., Hammerling, D. M., Levy, M. N., Xu, H., Dennis, J. M., Eaton, B. E., Edwards, J., Hannay, C., Mickelson, S. A., Neale, R. B., Nychka, D., Shollenberger, J., Tribbia, J., Vertenstein, M., and Williamson, D.: A new ensemble-based consistency test for the Community Earth System Model (pyCECT v1.0), Geosci. Model Dev., 8, 2829–2840, <a href="https://doi.org/10.5194/gmd-8-2829-2015" target="_blank">https://doi.org/10.5194/gmd-8-2829-2015</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Benjamini and Hochberg(1995)</label><mixed-citation>
      
Benjamini, Y. and Hochberg, Y.: Controlling the False Discovery Rate: A
Practical and Powerful Approach to Multiple Testing, Journal of the Royal
Statistical Society: Series B (Methodological), 57, 289–300,
<a href="https://doi.org/10.1111/j.2517-6161.1995.tb02031.x" target="_blank">https://doi.org/10.1111/j.2517-6161.1995.tb02031.x</a>, 1995.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Benjamini and Yekutieli(2001)</label><mixed-citation>
      
Benjamini, Y. and Yekutieli, D.: The control of the false discovery rate in
multiple testing under dependency, The Annals of Statistics, 29, 1165–1188, <a href="https://doi.org/10.1214/aos/1013699998" target="_blank">https://doi.org/10.1214/aos/1013699998</a>, 2001.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Bi et al.(2022)</label><mixed-citation>
      
Bi, K.,  Xie, L.,  Zhang, H., Chen, X., Gu, X., and Tian, Q.: Accurate medium-range global weather
forecasting with 3D neural networks, Nature, 610, 87–93, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Burrell et al.(2020)</label><mixed-citation>
      
Burrell, A. L., Evans, J. P., and De Kauwe, M. G.: Anthropogenic climate change
has driven over 5 million km<sup>2</sup> of drylands towards desertification, Nature
Communications, 11, 3853, <a href="https://doi.org/10.1038/s41467-020-17710-7" target="_blank">https://doi.org/10.1038/s41467-020-17710-7</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Büttner et al.(2024)</label><mixed-citation>
      
Büttner, M., Alt, C., Kenter, T., Köstler, H., Plessl, C., and
Aizinger, V.: Enabling Performance Portability for Shallow Water Equations on
CPUs, GPUs, and FPGAs with SYCL, in: Proceedings of the Platform for Advanced
Scientific Computing Conference, PASC '24, Association for Computing
Machinery, New York, NY, USA, ISBN 9798400706394,
<a href="https://doi.org/10.1145/3659914.3659925" target="_blank">https://doi.org/10.1145/3659914.3659925</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Dask Development Team(2016)</label><mixed-citation>
      
Dask Development Team: Dask: Library for dynamic task scheduling,
<a href="http://dask.pydata.org" target="_blank"/> (last access: 24 October 2025), 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Dueben and Bauer(2018)</label><mixed-citation>
      
Dueben, P. D. and Bauer, P.: Challenges and design choices for global weather and climate models based on machine learning, Geosci. Model Dev., 11, 3999–4009, <a href="https://doi.org/10.5194/gmd-11-3999-2018" target="_blank">https://doi.org/10.5194/gmd-11-3999-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>E3SM Project(2023)</label><mixed-citation>
      
E3SM Project, D.: Energy Exascale Earth System Model v2.1.0, DOE Code [software],  <a href="https://doi.org/10.11578/E3SM/dc.20230110.5" target="_blank">https://doi.org/10.11578/E3SM/dc.20230110.5</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Eidhammer et al.(2024)</label><mixed-citation>
      
Eidhammer, T., Gettelman, A., Thayer-Calder, K., Watson-Parris, D., Elsaesser, G., Morrison, H., van Lier-Walqui, M., Song, C., and McCoy, D.: An extensible perturbed parameter ensemble for the Community Atmosphere Model version 6, Geosci. Model Dev., 17, 7835–7853, <a href="https://doi.org/10.5194/gmd-17-7835-2024" target="_blank">https://doi.org/10.5194/gmd-17-7835-2024</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Gentle(2003)</label><mixed-citation>
      
Gentle, J. E.: Random Number Generation and Monte Carlo Methods,
Springer-Verlag, ISBN 0387001786, <a href="https://doi.org/10.1007/b97336" target="_blank">https://doi.org/10.1007/b97336</a>, 2003.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Golaz et al.(2022)</label><mixed-citation>
      
Golaz, J.-C., Van Roekel, L. P., Zheng, X., Roberts, A. F., Wolfe, J. D., Lin,
W., Bradley, A. M., Tang, Q., Maltrud, M. E., Forsyth, R. M., Zhang, C.,
Zhou, T., Zhang, K., Zender, C. S., Wu, M., Wang, H., Turner, A. K., Singh,
B., Richter, J. H., Qin, Y., Petersen, M. R., Mametjanov, A., Ma, P.-L.,
Larson, V. E., Krishna, J., Keen, N. D., Jeffery, N., Hunke, E. C., Hannah,
W. M., Guba, O., Griffin, B. M., Feng, Y., Engwirda, D., Di Vittorio, A. V.,
Dang, C., Conlon, L. M., Chen, C.-C.-J., Brunke, M. A., Bisht, G., Benedict,
J. J., Asay-Davis, X. S., Zhang, Y., Zhang, M., Zeng, X., Xie, S., Wolfram,
P. J., Vo, T., Veneziani, M., Tesfa, T. K., Sreepathi, S., Salinger, A. G.,
Reeves Eyre, J. E. J., Prather, M. J., Mahajan, S., Li, Q., Jones, P. W.,
Jacob, R. L., Huebler, G. W., Huang, X., Hillman, B. R., Harrop, B. E.,
Foucar, J. G., Fang, Y., Comeau, D. S., Caldwell, P. M., Bartoletti, T.,
Balaguru, K., Taylor, M. A., McCoy, R. B., Leung, L. R., and Bader, D. C.:
The DOE E3SM Model Version 2: Overview of the Physical Model and Initial
Model Evaluation, Journal of Advances in Modeling Earth Systems, 14,
e2022MS003156, <a href="https://doi.org/10.1029/2022MS003156" target="_blank">https://doi.org/10.1029/2022MS003156</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Harris et al.(2020)</label><mixed-citation>
      
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P.,
Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R.,
Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del
Río, J. F., Wiebe, M., Peterson, P., Gérard-Marchant, P.,
Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant,
T. E.: Array programming with NumPy, Nature, 585, 357–362,
<a href="https://doi.org/10.1038/s41586-020-2649-2" target="_blank">https://doi.org/10.1038/s41586-020-2649-2</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Hoyer and Hamman(2017)</label><mixed-citation>
      
Hoyer, S. and Hamman, J.: xarray: N-D labeled arrays and datasets in
Python, Journal of Open Research Software, 5, <a href="https://doi.org/10.5334/jors.148" target="_blank">https://doi.org/10.5334/jors.148</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Hunter(2007)</label><mixed-citation>
      
Hunter, J. D.: Matplotlib: A 2D graphics environment, Computing in Science &amp;
Engineering, 9, 90–95, <a href="https://doi.org/10.1109/MCSE.2007.55" target="_blank">https://doi.org/10.1109/MCSE.2007.55</a>, 2007.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Intel Corporation(2023)</label><mixed-citation>
      
Intel Corporation: Intel Fortran Compiler Developer Guide and Reference,
<a href="https://www.intel.com/content/www/us/en/docs/fortran-compiler/developer-guide-reference/2023-0/compiler-options-001.html" target="_blank"/> (last access: 1 May 2025),
2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Kelleher and Mahajan(2025a)</label><mixed-citation>
      
Kelleher, M. and Mahajan, S.: Detectable Climate (v1.1.0), Zenodo [code], <a href="https://doi.org/10.5281/zenodo.17438094" target="_blank">https://doi.org/10.5281/zenodo.17438094</a>, 2025a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Kelleher and Mahajan(2025b)</label><mixed-citation>
      
Kelleher, M. and Mahajan, S.: Detectable Climate Bootstrap Data (Version v2), Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.17438071" target="_blank">https://doi.org/10.5281/zenodo.17438071</a>, 2025b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Mahajan(2021)</label><mixed-citation>
      
Mahajan, S.: Ensuring statistical reproducibility of ocean model simulations in
the age of hybrid computing, in: Proceedings of the Platform for Advanced
Scientific Computing Conference, PASC '21, Association for Computing
Machinery, New York, NY, USA, ISBN 9781450385633,
<a href="https://doi.org/10.1145/3468267.3470572" target="_blank">https://doi.org/10.1145/3468267.3470572</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Mahajan et al.(2017)</label><mixed-citation>
      
Mahajan, S., Gaddis, A. L., Evans, K. J., and Norman, M. R.: Exploring an
Ensemble-Based Approach to Atmospheric Climate Modeling and Testing at Scale, international Conference
on Computational Science, ICCS 2017, 12-14 June 2017, Zurich, Switzerland,
Procedia Computer Science, 108, 735–744,
<a href="https://doi.org/10.1016/j.procs.2017.05.259" target="_blank">https://doi.org/10.1016/j.procs.2017.05.259</a>,
2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Mahajan et al.(2019a)</label><mixed-citation>
      
Mahajan, S., Evans, K. J., Kennedy, J. H., Xu, M., and Norman, M. R.: A
multivariate approach to ensure statistical reproducibility of climate model
simulations, in: Proceedings of the Platform for Advanced Scientific
Computing Conference,  1–10, 2019a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Mahajan et al.(2019b)</label><mixed-citation>
      
Mahajan, S., Evans, K. J., Kennedy, J. H., Xu, M., Norman, M. R., and
Branstetter, M. L.: Ongoing solution reproducibility of earth system models
as they progress toward exascale computing, The International Journal of High
Performance Computing Applications, 33, 784–790,
<a href="https://doi.org/10.1177/1094342019837341" target="_blank">https://doi.org/10.1177/1094342019837341</a>, 2019b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Mahajan et al.(2022)</label><mixed-citation>
      
Mahajan, S., Tang, Q., Keen, N. D., Golaz, J.-C., and van Roekel, L. P.:
Simulation of ENSO teleconnections to precipitation extremes over the United
States in the high-resolution version of E3SM, Journal of Climate, 35,
3371–3393, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Mann and Whitney(1947)</label><mixed-citation>
      
Mann, H. and Whitney, D. R.: On a Test of Whether one of Two Random Variables
is Stochastically Larger than the Other., Ann. Math. Statist., 18, 50–60,
<a href="https://doi.org/10.1214/aoms/1177730491" target="_blank">https://doi.org/10.1214/aoms/1177730491</a>, 1947.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>McKinney(2010)</label><mixed-citation>
      
McKinney, W.: Data Structures for Statistical Computing in
Python, in: Proceedings of the 9th Python in Science Conference,
edited by: van der Walt, S. and Millman, J.,  56–61,
<a href="https://doi.org/10.25080/Majora-92bf1922-00a" target="_blank">https://doi.org/10.25080/Majora-92bf1922-00a</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Mielikainen et al.(2016)</label><mixed-citation>
      
Mielikainen, J., Price, E., Huang, B., Huang, H.-L. A., and Lee, T.: GPU
Compute Unified Device Architecture (CUDA)-based Parallelization of the RRTMG
Shortwave Rapid Radiative Transfer Model, IEEE Journal of Selected Topics in
Applied Earth Observations and Remote Sensing, 9, 921–931,
<a href="https://doi.org/10.1109/JSTARS.2015.2427652" target="_blank">https://doi.org/10.1109/JSTARS.2015.2427652</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Milroy et al.(2018)</label><mixed-citation>
      
Milroy, D. J., Baker, A. H., Hammerling, D. M., and Jessup, E. R.: Nine time steps: ultra-fast statistical consistency testing of the Community Earth System Model (pyCECT v3.0), Geosci. Model Dev., 11, 697–711, <a href="https://doi.org/10.5194/gmd-11-697-2018" target="_blank">https://doi.org/10.5194/gmd-11-697-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Pathak et al.(2022)</label><mixed-citation>
      
Pathak, J., Subramanian, S., Harrington, P., Raja, S., Chattopadhyay, A.,  Mardani, M.,  Kurth, T.,  Hall, D.,  Li, Z.,  Azizzadenesheli, K.,  Hassanzadeh, P.,  Kashinath, K., and Anandkumar, A.: FourCastNet: A Global Data-driven High-resolution Weather Model using Adaptive Fourier Neural Operators, arXiv [preprint], <a href="https://doi.org/10.48550/arXiv.2202.11214" target="_blank">https://doi.org/10.48550/arXiv.2202.11214</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Qian et al.(2018)</label><mixed-citation>
      
Qian, Y., Wan, H., Yang, B., Golaz, J.-C., Harrop, B., Hou, Z., Larson, V. E.,
Leung, L. R., Lin, G., Lin, W., Ma, P.-L., Ma, H.-Y., Rasch, P., Singh, B.,
Wang, H., Xie, S., and Zhang, K.: Parametric Sensitivity and Uncertainty
Quantification in the Version 1 of E3SM Atmosphere Model Based on
Short Perturbed Parameter Ensemble Simulations, Journal of
Geophysical Research: Atmospheres, 123, 13046–13073,
<a href="https://doi.org/10.1029/2018JD028927" target="_blank">https://doi.org/10.1029/2018JD028927</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Rasp et al.(2020)</label><mixed-citation>
      
Rasp, S., Dueben, P. D., Scher, S., Weyn, J. A., Mouatadid, S., and Thuerey,
N.: WeatherBench: A benchmark dataset for data-driven weather forecasting,
Journal of Advances in Modeling Earth Systems, 12, e2020MS002203, <a href="https://doi.org/10.1029/2020MS002203" target="_blank">https://doi.org/10.1029/2020MS002203</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Renard et al.(2008)</label><mixed-citation>
      
Renard, B., Lang, M., Bois, P., Dupeyrat, A., Mestre, O., Niel, H.,  Sauquet, E., Prudhomme, C., Parey, S., Paquet, E., Neppel, L., and Gailhard, J.: Regional methods for trend
detection: Assessing field significance and regional consistency, Water
Resources Research, 44, <a href="https://doi.org/10.1029/2007WR006268" target="_blank">https://doi.org/10.1029/2007WR006268</a>, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Rosinski and Williamson(1997)</label><mixed-citation>
      
Rosinski, J. M. and Williamson, D. L.: The Accumulation of Rounding Errors and
Port Validation for Global Atmospheric Models, SIAM Journal on Scientific
Computing, 18, 552–564, <a href="https://doi.org/10.1137/S1064827594275534" target="_blank">https://doi.org/10.1137/S1064827594275534</a>, 1997.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Seabold and Perktold(2010)</label><mixed-citation>
      
Seabold, S. and Perktold, J.: statsmodels: Econometric and statistical modeling
with Python, in: 9th Python in Science Conference, 28 June–3 July 2010, Austin, TX, USA,
<a href="https://doi.org/10.25080/Majora-92bf1922-012" target="_blank">https://doi.org/10.25080/Majora-92bf1922-012</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Ventura et al.(2004)</label><mixed-citation>
      
Ventura, V., Paciorek, C. J., and Risbey, J. S.: Controlling the Proportion of
Falsely Rejected Hypotheses when Conducting Multiple Tests with
Climatological Data, Journal of Climate, 17, 4343–4356,
<a href="https://doi.org/10.1175/3199.1" target="_blank">https://doi.org/10.1175/3199.1</a>, 2004.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Virtanen et al.(2020)</label><mixed-citation>
      
Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T.,
Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van
der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson,
A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J., Polat, İ., Feng,
Y., Moore, E. W., VanderPlas, J., Laxalde, D., Perktold, J., Cimrman, R.,
Henriksen, I., Quintero, E. A., Harris, C. R., Archibald, A. M., Ribeiro,
A. H., Pedregosa, F., van Mulbregt, P., and SciPy 1.0 Contributors:
SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python,
Nature Methods, 17, 261–272, <a href="https://doi.org/10.1038/s41592-019-0686-2" target="_blank">https://doi.org/10.1038/s41592-019-0686-2</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Wan et al.(2017)</label><mixed-citation>
      
Wan, H., Zhang, K., Rasch, P. J., Singh, B., Chen, X., and Edwards, J.: A new and inexpensive non-bit-for-bit solution reproducibility test based on time step convergence (TSC1.0), Geosci. Model Dev., 10, 537–552, <a href="https://doi.org/10.5194/gmd-10-537-2017" target="_blank">https://doi.org/10.5194/gmd-10-537-2017</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Waskom(2021)</label><mixed-citation>
      
Waskom, M. L.: seaborn: statistical data visualization, Journal of Open Source
Software, 6, 3021, <a href="https://doi.org/10.21105/joss.03021" target="_blank">https://doi.org/10.21105/joss.03021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Watson-Parris et al.(2022)</label><mixed-citation>
      
Watson-Parris, D., Rao, Y., Olivié, D., Seland, Ø., Nowack, P., Camps-Valls, G., Stier, P., Bouabid, S., Dewey, M., Fons, E., Gonzalez, J., Harder, P., Jeggle, K., Lenhardt, J., Manshausen, P., Novitasari, M., Ricard, L., and Roesch, C.: ClimateBench v1.0: A Benchmark for Data-Driven Climate Projections, Journal of Advances in Modeling Earth Systems, 14,  <a href="https://doi.org/10.1029/2021MS002954" target="_blank">https://doi.org/10.1029/2021MS002954</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Whan and Zwiers(2017)</label><mixed-citation>
      
Whan, K. and Zwiers, F.: The impact of ENSO and the NAO on extreme winter
precipitation in North America in observations and regional climate models,
Climate Dynamics, 48, 1401–1411, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Wilks(2006)</label><mixed-citation>
      
Wilks, D. S.: On “Field Significance” and the False Discovery Rate, Journal
of Applied Meteorology and Climatology, 45, 1181–1189,
<a href="https://doi.org/10.1175/JAM2404.1" target="_blank">https://doi.org/10.1175/JAM2404.1</a>, 2006.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Wilks(2016)</label><mixed-citation>
      
Wilks, D. S.: “The Stippling Shows Statistically Significant Grid Points”:
How Research Results are Routinely Overstated and Overinterpreted, and What
to Do about It, Bulletin of the American Meteorological Society, 97, 2263–2273, <a href="https://doi.org/10.1175/BAMS-D-15-00267.1" target="_blank">https://doi.org/10.1175/BAMS-D-15-00267.1</a>, 2016.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>Zeman and Schär(2022)</label><mixed-citation>
      
Zeman, C. and Schär, C.: An ensemble-based statistical methodology to detect differences in weather and climate model executables, Geosci. Model Dev., 15, 3183–3203, <a href="https://doi.org/10.5194/gmd-15-3183-2022" target="_blank">https://doi.org/10.5194/gmd-15-3183-2022</a>, 2022.

    </mixed-citation></ref-html>--></article>
