<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">ESD</journal-id><journal-title-group>
    <journal-title>Earth System Dynamics</journal-title>
    <abbrev-journal-title abbrev-type="publisher">ESD</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Earth Syst. Dynam.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">2190-4987</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/esd-11-885-2020</article-id><title-group><article-title>How large does a large ensemble need to be?</article-title><alt-title>How large does a large ensemble need to be?</alt-title>
      </title-group><?xmltex \runningtitle{How large does a large ensemble need to be?}?><?xmltex \runningauthor{S.~Milinski et al.}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <name><surname>Milinski</surname><given-names>Sebastian</given-names></name>
          <email>sebastian.milinski@mpimet.mpg.de</email>
        <ext-link>https://orcid.org/0000-0001-5216-4662</ext-link></contrib>
        <contrib contrib-type="author" corresp="no">
          <name><surname>Maher</surname><given-names>Nicola</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-3922-9833</ext-link></contrib>
        <contrib contrib-type="author" corresp="no">
          <name><surname>Olonscheck</surname><given-names>Dirk</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-8962-1406</ext-link></contrib>
        <aff id="aff1"><institution>Max Planck Institute for Meteorology, Hamburg, Germany</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Sebastian Milinski (sebastian.milinski@mpimet.mpg.de)</corresp></author-notes><pub-date><day>30</day><month>October</month><year>2020</year></pub-date>
      
      <volume>11</volume>
      <issue>4</issue>
      <fpage>885</fpage><lpage>901</lpage>
      <history>
        <date date-type="received"><day>11</day><month>November</month><year>2019</year></date>
           <date date-type="rev-request"><day>29</day><month>November</month><year>2019</year></date>
           <date date-type="rev-recd"><day>25</day><month>August</month><year>2020</year></date>
           <date date-type="accepted"><day>7</day><month>September</month><year>2020</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2020 Sebastian Milinski et al.</copyright-statement>
        <copyright-year>2020</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://esd.copernicus.org/articles/11/885/2020/esd-11-885-2020.html">This article is available from https://esd.copernicus.org/articles/11/885/2020/esd-11-885-2020.html</self-uri><self-uri xlink:href="https://esd.copernicus.org/articles/11/885/2020/esd-11-885-2020.pdf">The full text article is available as a PDF file from https://esd.copernicus.org/articles/11/885/2020/esd-11-885-2020.pdf</self-uri>
      <abstract><title>Abstract</title>
    <p id="d1e95">Initial-condition large ensembles with ensemble sizes ranging from 30 to 100 members have become a commonly used tool for quantifying the forced response and internal variability in various components of the climate system. However, there is no consensus on the ideal or even sufficient ensemble size for a large ensemble. Here, we introduce an objective method to estimate the required ensemble size that can be applied to any given application and demonstrate its use on the examples of global mean near-surface air temperature, local temperature and precipitation, and variability in the El Niño–Southern Oscillation (ENSO) region and central United States for the Max Planck Institute Grand Ensemble (MPI-GE). Estimating the required ensemble size is relevant not only for designing or choosing a large ensemble but also for designing targeted sensitivity experiments with a model. Where possible, we base our estimate of the required ensemble size on the pre-industrial control simulation, which is available for every model. We show that more ensemble members are needed to quantify variability than the forced response, with the largest ensemble sizes needed to detect changes in internal variability itself. Finally, we highlight that the required ensemble size depends on both the acceptable error to the user and the studied quantity.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d1e107">Single model initial-condition large ensembles (SMILEs) are a valuable tool for cleanly separating a model's forced response from internal variability and improving our understanding of the observed trajectory of the climate system in the past, as well as its projected future evolution <xref ref-type="bibr" rid="bib1.bibx32 bib1.bibx7 bib1.bibx24 bib1.bibx17 bib1.bibx21 bib1.bibx3 bib1.bibx30 bib1.bibx18 bib1.bibx14 bib1.bibx27" id="paren.1"/>.</p>
      <p id="d1e113">The ensemble sizes currently available for individual global coupled climate models greatly differ. The single-model ensembles within the Coupled Model Intercomparison Project Phase 5 and 6 (CMIP5, CMIP6) are on the low end of available ensemble sizes, typically ranging from 3 to 10 ensemble members for a model, with the majority of models having only one member available. In contrast, computationally expensive SMILEs position themselves on the top end of available ensemble sizes, providing up to 200 ensemble members for a single model and forcing scenario. While studies are beginning to compare multiple SMILEs <xref ref-type="bibr" rid="bib1.bibx20 bib1.bibx9" id="paren.2"/>, there is still no clear consensus on how large such an ensemble should be for any given application.</p>
      <p id="d1e119">We here introduce a new framework to objectively estimate the required ensemble size for different types of questions and make use of a model's pre-industrial control simulation where possible. Using the pre-industrial control simulation allows us to estimate the required ensemble size for a specific model even if no large ensemble is available. The objective approach can also help to allocate resources more efficiently <xref ref-type="bibr" rid="bib1.bibx11" id="paren.3"/> and to inform the modelling community how many ensemble members are desirable for CMIP models.</p>
      <p id="d1e125">One of the most common applications of SMILEs is to separate a forced response due to anthropogenic global warming from the noise of internal variability. In a sufficiently large ensemble the ensemble mean can be used as an estimator for the forced response <xref ref-type="bibr" rid="bib1.bibx13" id="paren.4"/>. This approach has been applied to study various regions and quantities.</p>
      <?pagebreak page886?><p id="d1e132">On a global scale, <xref ref-type="bibr" rid="bib1.bibx8" id="text.5"/> investigate the forced response in temperature and precipitation. They found that around 10 ensemble members are sufficient to detect changes in the global mean land temperature in the next decade, while more than 40 ensemble members are required to detect changes in precipitation. When going further into the future when the signal becomes larger, they find that fewer members are sufficient to detect a forced change. If the signal is large enough, a single ensemble member is sufficient to detect a significant change compared to present-day conditions. This happens when the trajectory of the single member emerges from the range of internal variability for present-day conditions.</p>
      <p id="d1e138">On both global and regional scales, <xref ref-type="bibr" rid="bib1.bibx22" id="text.6"/> used both the CMIP5 multi-model ensemble and the Max Planck Institute Grand Ensemble (MPI-GE) to conclude that multiple small ensembles from different models are useful for quantifying the response uncertainty across different models.</p>
      <p id="d1e144">While a forced response in global mean temperature only requires a relatively small ensembles size, forced changes on a smaller regional scale can be more difficult to detect because of the larger variability. <xref ref-type="bibr" rid="bib1.bibx19" id="text.7"/> investigated the ocean carbon sink and found that up to 79 ensemble members are required to isolate a forced decadal trend in the RCP4.5 scenario in the Southern Ocean, a region with large internal variability. <xref ref-type="bibr" rid="bib1.bibx25" id="text.8"/> quantify the forced response in North Atlantic temperature and argue that for this region more than four ensemble members are required for a robust estimate of the forced response from a SMILE. Although the objective of the two studies is similar – identifying a forced response – the required ensemble size is very different, indicating that different regions and quantities can have very different requirements for the ensemble size.</p>
      <p id="d1e153">In addition to investigating forced changes to anthropogenic forcing, large ensembles also allow an investigation of forced responses to other external forcings such as volcanic eruptions. For regional temperature changes, <xref ref-type="bibr" rid="bib1.bibx23" id="text.9"/> find that up to 40 ensemble members are necessary for a robust detection of a temperature response after a volcanic eruption. <xref ref-type="bibr" rid="bib1.bibx2" id="text.10"/> investigate changes in atmospheric circulation after a volcanic eruption. They analyse the polar vortex and find that the required ensemble size to detect changes in the zonal wind after a strong volcanic eruption depends on the latitude: seven members are sufficient at the southward flank of the maximum positive wind anomaly, but up to 40 members are necessary to identify a response at high northern latitudes. The ratio of the signal to the noise from internal variability is different in different regions because the signal and/or the internal variability may differ. The target of <xref ref-type="bibr" rid="bib1.bibx2" id="text.11"/> was to detect a change in the circulation that is different from zero, but not to quantify it. Quantifying the magnitude of the forced response may require an even larger ensemble size for this application.</p>
      <p id="d1e165">Large ensembles have also been used to quantify internal variability, with some studies arguing that very large ensemble sizes are necessary: <xref ref-type="bibr" rid="bib1.bibx6" id="text.12"/> conclude that an ensemble with several hundred members is required to characterise a model's climate, while <xref ref-type="bibr" rid="bib1.bibx10" id="text.13"/> demonstrate that 100 members are sufficient.
On the other hand, some studies argue that the pre-industrial control simulation is sufficient to quantify internal variability and no large ensemble is required. <xref ref-type="bibr" rid="bib1.bibx29" id="text.14"/> argue that the pre-industrial control simulation can be used to provide a robust estimate of internal variability and represent future internal variability, implying that a single ensemble member for each model may be sufficient. However, this approach only works if the internal variability does not change over time. In addition, a single realisation for a transient scenario does not allow a clean separation of the forced response and internal variability, even if the magnitude of the internal variability is quantified using a pre-industrial control simulation.</p>
      <p id="d1e177">El Niño–Southern Oscillation (ENSO) variability and its potential changes under global warming have been investigated in several studies, and widely different future changes have been identified <xref ref-type="bibr" rid="bib1.bibx26 bib1.bibx1 bib1.bibx5" id="paren.15"/>.
<xref ref-type="bibr" rid="bib1.bibx20" id="text.16"/> investigate ENSO variability and its potential changes under global warming in several large ensembles. They find that at least 30 ensemble members are required for a robust estimate of ENSO variability. When using a smaller ensemble, sampling uncertainty may lead to false detection of a forced change in ENSO or a robust difference between two models.</p>
      <p id="d1e187">All of the aforementioned studies demonstrate that different applications require different ensemble sizes. However, these studies suffer from two drawbacks. First, the required ensemble size can only be estimated once a signal has been identified in a large ensemble, which requires the large ensemble to exist and be large enough in the first place. Second, the result might be model dependent and may only provide a very rough estimate of the required ensemble size when addressing the same question with a different model.</p>
      <p id="d1e190">In this paper, we introduce a basic recipe for estimating the required ensemble size in Sect. <xref ref-type="sec" rid="Ch1.S3"/>. The required or ideal ensemble size depends on the region and quantity that is investigated and the type of question. Therefore we differentiate three types of questions that represent questions typically addressed with large ensembles:
<list list-type="order"><list-item>
      <p id="d1e197">How many ensemble members are required to identify the response to a change in the external forcing? (Sect. <xref ref-type="sec" rid="Ch1.S4.SS1"/>)</p></list-item><list-item>
      <p id="d1e203">How many ensemble members are required to adequately sample the spectrum of internal variability? (Sect. <xref ref-type="sec" rid="Ch1.S4.SS2"/>)</p></list-item><list-item>
      <p id="d1e209">How many ensemble members are required to identify a forced change in internal variability (e.g., a mode of variability such as ENSO)? (Sect. <xref ref-type="sec" rid="Ch1.S4.SS3"/>)</p></list-item></list>
<?xmltex \hack{\newpage}?><?xmltex \hack{\noindent}?>An additional discussion of caveats associated with the choice of sampling method is discussed in Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/> and is relevant for users of the approach proposed in this study.</p>
</sec>
<?pagebreak page887?><sec id="Ch1.S2">
  <label>2</label><title>Model</title>
      <p id="d1e228">In this study, we are using simulations from the Max Planck Institute Grand Ensemble. The MPI-GE consists of large initial-condition ensembles for several experiments with the Max Planck Institute Earth System Model (MPI-ESM) in its low-resolution configuration. Ensemble members are generated by sampling different years from a 2000-year pre-industrial control simulation for the initial conditions (macro-initialisation). The forcing for the experiments follows the protocol of the CMIP5 simulations <xref ref-type="bibr" rid="bib1.bibx28" id="paren.17"/>. The model configuration and experiments are described in more detail in <xref ref-type="bibr" rid="bib1.bibx21" id="text.18"/>.</p>
      <p id="d1e237">In this study, we use three experiments from the MPI-GE:
<list list-type="bullet"><list-item>
      <p id="d1e242">pre-industrial control simulation (2000 years)</p></list-item><list-item>
      <p id="d1e246">historical simulations (1850–2005, 200 members)</p></list-item><list-item>
      <p id="d1e250">1 % <inline-formula><mml:math id="M1" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">CO</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> simulations (156 years, 100 members).</p></list-item></list></p>
      <p id="d1e264">Note that only the first 100 historical realisations are described in <xref ref-type="bibr" rid="bib1.bibx21" id="text.19"/>. Realisations 101–200 were added later and use the same configuration as the first 100 realisations but are initialised from different years of the pre-industrial control simulation.</p>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>A simple method to estimate the required ensemble size</title>
      <p id="d1e278">In this section, we use a simple example to design a generic recipe for estimating the required ensemble size for any given application. In Sect. <xref ref-type="sec" rid="Ch1.S4.SS1"/> to <xref ref-type="sec" rid="Ch1.S4.SS3"/>, we then apply this recipe to various examples.</p>
      <p id="d1e285">One of the most common applications of a large ensemble is the separation of the forced response and the random internal variability in a time series. Each realisation from a large ensemble is subject to the same external forcing. Due to different initial conditions, each realisation is a combination of the forced response due to this external forcing and a unique trajectory of quasi-random internal variability. By averaging over a large number of realisations, internal variability cancels out and the forced response remains <xref ref-type="bibr" rid="bib1.bibx12" id="paren.20"/>. Therefore, the ensemble mean of a large ensemble is often referred to as the forced response. Figure <xref ref-type="fig" rid="Ch1.F1"/> shows the ensemble mean global mean near-surface air temperature (GSAT, blue line) of 200 realisations with CMIP5 historical forcing from the MPI-GE <xref ref-type="bibr" rid="bib1.bibx21" id="paren.21"/>. Because of the large ensemble size and the use of a globally averaged quantity, the 200-member mean is a clean estimate of the forced response.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><?xmltex \currentcnt{1}?><label>Figure 1</label><caption><p id="d1e298">The forced response can be quantified using the ensemble mean in a large ensemble, while the ensemble mean of smaller ensembles still contains a contribution from internal variability. The figure is based on global and annual mean near-surface air temperature from the MPI-GE 200-member historical ensemble. The dark blue line shows the 200-member ensemble mean time series. Shaded regions show the range of forced responses estimated by resampling 1000 times for various ensemble sizes. The light grey shading shows the range of the full ensemble, i.e. the minimum to maximum of all 200 realisations for every single year.</p></caption>
        <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://esd.copernicus.org/articles/11/885/2020/esd-11-885-2020-f01.png"/>

      </fig>

      <p id="d1e308"><?xmltex \hack{\newpage}?>Assuming that the 200-member mean provides a good estimate of the forced response, we can then form subsets of the large ensemble to investigate how well the ensemble mean of a smaller ensemble can isolate the forced response. We draw 1000 random samples of sets of three members from MPI-GE without replacement. For each of these samples, the three-member ensemble mean is computed. The red envelope in Fig. <xref ref-type="fig" rid="Ch1.F1"/> shows the range of these 1000 samples of a three-member mean forced response. Compared to individual realisations (grey envelope), a three-member mean reduces internal variability, but it can deviate substantially from the 200-member mean. Repeating this analysis for 10, 20, and 50 members shows that a larger ensemble size can separate the forced response from internal variability more effectively.</p>
      <p id="d1e314">To quantify how effective the separation of forced response and internal variability is, we show the root-mean-square error (RMSE) of ensemble means for different ensemble sizes compared to the 200-member mean. The solid black line in Fig. <xref ref-type="fig" rid="Ch1.F2"/> shows how the expected RMSE decreases with increasing ensemble size until reaching zero for 200 members. By choosing an acceptable error, we can then determine the required ensemble size. For example, an acceptable error of 0.02 <inline-formula><mml:math id="M2" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>C would mean that an ensemble with approximately 50 members is required. We will return to the discussion of what constitutes an acceptable error in the examples in Sect. <xref ref-type="sec" rid="Ch1.S4.SS1"/> to <xref ref-type="sec" rid="Ch1.S4.SS3"/>.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><?xmltex \currentcnt{2}?><label>Figure 2</label><caption><p id="d1e334">A larger ensemble allows a more accurate quantification of the forced response. The black line shows the mean RMSE for GSAT for ensemble sizes from 2 to 200. The reference is the 200-member mean from Fig. <xref ref-type="fig" rid="Ch1.F1"/>, and the RMSE is computed for all 1000 samples. The shaded area shows the range of RMSE values for individual samples; the solid line shows the mean RMSE.</p></caption>
        <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://esd.copernicus.org/articles/11/885/2020/esd-11-885-2020-f02.png"/>

      </fig>

      <p id="d1e345">While a reduction in the error with increasing ensemble size is expected and indicates that a larger ensemble allows a more accurate representation of the forced response, the vanishing error when using 200 members occurs by construction because we assume that the 200-member mean represents the true forced response. How fast the error is converging therefore depends on how the random samples are generated.</p>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>A cautionary note on resampling</title>
      <p id="d1e355">One difficulty when determining the required ensemble size for a specific question is the chosen sampling approach: in this study, we generate synthetic ensembles of different ensemble sizes by randomly sampling members from a 200-member ensemble without replacement. Samples generated in this way are not fully independent when approaching the full ensemble size. For example, two random samples of 190 out of the available 200 members will share most of their members. This resampling introduces a problem when the signal is defined by using the full ensemble. Any subsample that is close to the full ensemble size will then indicate that the ensemble size is sufficient by construction.</p>
      <p id="d1e358">The resampling problem occurs with any limited sample. At some point, the 1000 random subsamples are not independent anymore because they share many of the randomly drawn members from the full ensemble. Therefore, they look more similar not only to each other but also to the 200-member mean. To demonstrate how this resampling affects our estimate of the error, we deliberately reduce the size of<?pagebreak page888?> the ensemble, for instance by only using the first 150 members and repeating the analysis. In an empirical analysis, we find that samples using more than 50 % of the available ensemble size to generate random samples lead to a substantial bias in the error estimate. We therefore recommend treating results indicating that, for example, more than 100 out of 200 members are required with caution because the true required ensemble size might be much larger. A more detailed discussion is provided in Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/>.</p>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>A recipe for estimating ensemble size</title>
      <p id="d1e372">Based on the example introduced in this section, we suggest the following approach to derive a robust estimate of the required ensemble size for any application. This method can be applied either to one of the existing large ensembles, as shown above for the MPI-GE, or to a long control run, which is available for all models participating in CMIP. We summarise the method in five steps before applying it to several examples in the next section:
<?xmltex \hack{\newpage}?>
<list list-type="order"><list-item>
      <p id="d1e379">Define the question to be addressed (isolate a forced response, quantify variability, or detect a change in variability).</p></list-item><list-item>
      <p id="d1e383">Choose an error metric (e.g. RMSE or variance across samples) and an upper threshold based on the maximum error that is acceptable in the specific application.</p></list-item><list-item>
      <p id="d1e387">Estimate the error for different ensemble sizes by subsampling a long control run or a large ensemble of transient simulations.</p></list-item><list-item>
      <p id="d1e391">Determine the minimum ensemble size that is required to reduce the error below the threshold chosen in step 2.</p></list-item><list-item>
      <p id="d1e395">If the ensemble size determined in this way is less than 50 % of the available sample size (e.g. 50 members when subsampling a 100-member ensemble), then the estimated required ensemble size provides a robust estimate for the specific question and model investigated. If the estimated required ensemble size is larger than 50 % of the available sample size, then the estimate is biased low and the true required ensemble size could be substantially larger.</p></list-item></list></p>
</sec>
</sec>
<?pagebreak page889?><sec id="Ch1.S4">
  <label>4</label><title>Estimating the required ensemble size: applications</title>
      <p id="d1e407">In this section we use the pre-industrial control simulation and transient forced simulations from  the MPI-GE to estimate the required ensemble size for a variety of applications, ranging from global to regional quantities. We investigate the different aspects of quantifying the forced response or quantifying internal variability.</p>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Quantifying the forced response</title>
      <p id="d1e417">The forced response shown in Fig. <xref ref-type="fig" rid="Ch1.F1"/> contains various signals. The most prominent signal is the long-term warming trend caused by anthropogenic greenhouse gas emissions. On shorter timescales, volcanic eruptions lead to a cooling of the global mean surface temperature.</p>
      <p id="d1e422">In the first example, we continue to use the RMSE to quantify how well the entire forced response is estimated, but we move from the global mean to the regional forced response in near-surface air temperature in the historical runs from the MPI-GE. In Fig. <xref ref-type="fig" rid="Ch1.F3"/>a–e, the expected RMSE for each grid point is shown for ensemble sizes of 3, 5, 10, 50, and 100 members. This analysis is equivalent to the computation of the mean RMSE for GSAT (black line in Fig. <xref ref-type="fig" rid="Ch1.F2"/>), but applied to each grid point separately. The RMSE is computed as the mean difference between 100 samples and the 200-member mean. When the ensemble mean is based on just three members, the expected error in the estimated forced response is large over land regions, in particular in the Northern Hemisphere. Over the ocean, the RMSE is already small in many regions. Increasing the ensemble size reduces the error. At 50 members, the error is small in most regions of the globe. Because 50 members is smaller than 50 % of the maximum ensemble size (200 members), the error estimate for this ensemble size is reliable.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F3" specific-use="star"><?xmltex \currentcnt{3}?><label>Figure 3</label><caption><p id="d1e431"><bold>(a–e)</bold> The mean RMSE for the forced response in historical monthly mean near-surface air temperature of MPI-GE for <bold>(a)</bold> 3, <bold>(b)</bold> 5, <bold>(c)</bold> 10, <bold>(d)</bold> 50, and <bold>(e)</bold> 100 ensemble members relative to the 200-member mean, globally. The RMSE shown here is the mean from 100 random samples without replacement. <bold>(f–j)</bold> Required ensemble size to capture the 200-member mean forced response in historical monthly mean near-surface air temperature dependent on the acceptable error of <bold>(f)</bold> 0.1, <bold>(g)</bold> 0.2, <bold>(h)</bold> 0.3, <bold>(i)</bold> 0.5, and <bold>(j)</bold> 1.0 <inline-formula><mml:math id="M3" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>C.</p></caption>
          <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://esd.copernicus.org/articles/11/885/2020/esd-11-885-2020-f03.png"/>

        </fig>

      <p id="d1e487">To estimate how many members are sufficient to reduce the error below a critical threshold, we first need to determine what is an acceptable error as outlined in step 2 of the recipe. This choice will depend on the region of interest and the accuracy with which the forced response needs to be quantified. In Fig. <xref ref-type="fig" rid="Ch1.F3"/>f–j, we show how many members are necessary to estimate the forced response in near-surface air temperature for five acceptable errors that were chosen for illustrative purposes. If the acceptable error (RMSE) is 0.1 <inline-formula><mml:math id="M4" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>C, 10–30 ensemble members are sufficient over the tropical ocean, while more than 50 ensemble members are required over most land regions. Beyond 100 members, the resampling problem inhibits reliable estimates of the sufficient ensemble size. For an acceptable error of 0.25 <inline-formula><mml:math id="M5" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>C, fewer than 10 members are sufficient over most ocean regions, while more than 50 members are required over high-northern-latitude land regions. For an acceptable error of 0.5 <inline-formula><mml:math id="M6" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula>C, only high-latitude land regions require a large ensemble, while the forced response over ocean and land regions at lower latitudes can be estimated with fewer than 10 members.</p>
      <p id="d1e519">Conversely for rainfall, the error in estimating the forced signal when using a small ensemble is larger over the tropics than over the higher latitudes (Fig. <xref ref-type="fig" rid="Ch1.F4"/>a–e). The largest errors can be found over the Indian Ocean and western tropical Pacific. Similar to temperature, a 50-member ensemble shows very small errors across the globe.</p>
      <p id="d1e524">In Fig. <xref ref-type="fig" rid="Ch1.F4"/>f–j we show how many members are necessary to estimate the forced response with an acceptable error of 0.1, 0.2, 0.3, 0.5, and 1 mm d<inline-formula><mml:math id="M7" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. For an acceptable error of 0.2 mm d<inline-formula><mml:math id="M8" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, some ocean regions require more than 100 members to capture the forced rainfall response with the required accuracy, while fewer than 20 members are sufficient over northern Africa and Eurasia. Over large parts of North America, between 20 and 40 members are required to estimate the forced rainfall response. For an acceptable error of 0.5 mm d<inline-formula><mml:math id="M9" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, 20 to 40 members are required over the Indian Ocean and western tropical Pacific, while fewer than 10 members are sufficient elsewhere.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F4" specific-use="star"><?xmltex \currentcnt{4}?><label>Figure 4</label><caption><p id="d1e567"><bold>(a–e)</bold> The mean RMSE for the forced response in historical monthly mean total precipitation of MPI-GE for <bold>(a)</bold> 3, <bold>(b)</bold> 5, <bold>(c)</bold> 10, <bold>(d)</bold> 50, and <bold>(e)</bold> 100 ensemble members relative to the 200-member mean, globally. The RMSE shown here is the mean from 100 random samples without replacement. <bold>(f–j)</bold> Required ensemble size to capture the 200-member mean forced response in historical monthly mean total precipitation dependent on the acceptable error of <bold>(f)</bold> 0.1, <bold>(g)</bold> 0.2, <bold>(h)</bold> 0.3, <bold>(i)</bold> 0.5, and <bold>(j)</bold> 1.0 mm d<inline-formula><mml:math id="M10" display="inline"><mml:msup><mml:mi/><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>.</p></caption>
          <?xmltex \igopts{width=497.923228pt}?><graphic xlink:href="https://esd.copernicus.org/articles/11/885/2020/esd-11-885-2020-f04.png"/>

        </fig>

      <p id="d1e625">For the example in Figs. <xref ref-type="fig" rid="Ch1.F3"/> and <xref ref-type="fig" rid="Ch1.F4"/>, the objective was to isolate the full forced response in a time series, defined as the 200-member ensemble mean time series at every grid point. The full forced response includes all external forcings, both natural and anthropogenic. In many applications, the objective might be to isolate a specific feature of the forced response rather than all components. In the following two examples, we will demonstrate how to estimate the required ensemble size needed to isolate the global warming trend in the 20th century and the global cooling after a major volcanic eruption.</p>
      <?pagebreak page890?><p id="d1e633">The global warming signal follows a much simpler trajectory than the forced response to all external forcings (cf. Fig. <xref ref-type="fig" rid="Ch1.F1"/>). Here, we fit a linear trend to the historical time series for 1920 to 2005 and define the 200-member mean as the true forced warming trend. Over the 68-year period from 1920 to 2005, the model warms by 0.65 K (Fig. <xref ref-type="fig" rid="Ch1.F5"/>). We acknowledge that a linear trend may not represent the anthropogenic warming accurately but use this definition to illustrate how a specific aspect of the forced response can be investigated.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F5" specific-use="star"><?xmltex \currentcnt{5}?><label>Figure 5</label><caption><p id="d1e642">Linear warming trend from 1920 to 2005 for different ensemble sizes shown as a linear trend fitted to the ensemble mean. Black lines show maximum and minimum 86-year ensemble mean temperature trend from 1000 random samples. Errors are shown as percentage of the 200-member ensemble mean temperature trend.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://esd.copernicus.org/articles/11/885/2020/esd-11-885-2020-f05.png"/>

        </fig>

      <p id="d1e651">We subsample the ensemble for smaller ensemble sizes to generate forced warming trends for smaller ensemble sizes. While the trends in a single realisation can be anywhere in the range from 0.4 K to more than 0.8 K warming over 68 years, increasing the ensemble size to even five members leads to a significant reduction in the error (Fig. <xref ref-type="fig" rid="Ch1.F5"/>). The warming trend in every 10-member ensemble is within the 20 % range (<inline-formula><mml:math id="M11" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula> %, cyan dashed lines) of the true warming trend, indicating that ensembles with 5–10 members can provide a good estimate of the forced linear warming trend. While an error within the 20 % range of the true signal may be sufficient for some applications, the acceptable error for other applications might be larger or smaller and result in a smaller or larger acceptable ensemble size. For an acceptable error of <inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">15</mml:mn></mml:mrow></mml:math></inline-formula> %, five ensemble members would be sufficient, while for an acceptable error of <inline-formula><mml:math id="M13" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> % at least 25 ensemble members are required. All of these error estimates are below 100 members and therefore not dominated by the resampling problem.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F6" specific-use="star"><?xmltex \currentcnt{6}?><label>Figure 6</label><caption><p id="d1e688">GSAT cooling after Krakatoa eruption for different ensemble sizes shown as the ensemble mean temperature difference between 1882 and 1884. Black lines show maximum and minimum temperature response from 1000 random samples. Errors are shown as percentage of the 200-member ensemble mean temperature response.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://esd.copernicus.org/articles/11/885/2020/esd-11-885-2020-f06.png"/>

        </fig>

      <?pagebreak page891?><p id="d1e697">For signals on shorter timescales, the required ensemble size can be quite different. In Fig. <xref ref-type="fig" rid="Ch1.F6"/> we analyse the GSAT cooling after the Krakatoa eruption in 1883. The forced cooling is quantified as  the difference between 1884, the year after the eruption, and 1882, the year before the eruption. The 200-member mean shows a forced cooling of <inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.34</mml:mn></mml:mrow></mml:math></inline-formula> K after the eruption. Due to internal variability, a single realisation can even show a warming after the volcanic eruption. More than one member is required for the ensemble mean to capture a cooling in all samples. However, the ensemble mean cooling for five members can still exceed the range from <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.2</mml:mn></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn></mml:mrow></mml:math></inline-formula> K. More than 50 ensemble members are necessary to estimate the forced cooling within <inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">15</mml:mn></mml:mrow></mml:math></inline-formula> % of the true forced cooling, and approximately 100 members are required to reduce the error below <inline-formula><mml:math id="M18" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula> %. Due to the resampling problem, we cannot derive a robust estimate for the ensemble size required to reduce the error to less than <inline-formula><mml:math id="M19" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> %. While the analysis in Fig. <xref ref-type="fig" rid="Ch1.F6"/> suggests that 150 members would be sufficient for a <inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> % error, this number is close to the full ensemble size of 200 members and therefore biased low. The true required ensemble size to reduce the error to <inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> % is likely larger than 150 members.</p>
      <p id="d1e786">These examples demonstrate that the required sample size to estimate the forced response depends on the region and variable (Figs. <xref ref-type="fig" rid="Ch1.F3"/> and <xref ref-type="fig" rid="Ch1.F4"/>), as well as the feature of interest in the forced response (Figs. <xref ref-type="fig" rid="Ch1.F5"/> and <xref ref-type="fig" rid="Ch1.F6"/>). Whereas for some applications five members are sufficient to reduce the error to an acceptable magnitude, other applications require at least 50 members. A robust estimate for the forced response is given by the ensemble mean when averaging over the ensemble attenuates internal variability sufficiently <xref ref-type="bibr" rid="bib1.bibx13" id="paren.22"/>. The number of members required for this depends both on the magnitude of the forced signal and the magnitude of internal variability, as well as on the acceptable error for a specific application.</p>
</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Quantifying internal variability</title>
      <p id="d1e808">While quantifying the forced response only requires a robust estimate of the mean, quantifying internal variability requires<?pagebreak page892?> more members because higher-order moments of the distribution need to be estimated. In the following two examples, we use the second statistical moment of the distribution, the standard deviation, to quantify internal variability. We note that, if the distribution deviates from a normal distribution, only using the standard deviation to quantify internal variability may not be sufficient.</p>
      <p id="d1e811">Here, we investigate internal variability in two regions: the tropical Pacific, where the variability is primarily driven by ENSO, and the central United States (34–46<inline-formula><mml:math id="M22" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N, 116–96<inline-formula><mml:math id="M23" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> W). The tropical Pacific region shows substantial variability on interannual to decadal timescales. Previous work has demonstrated that large sample sizes are necessary to quantify ENSO variability <xref ref-type="bibr" rid="bib1.bibx20 bib1.bibx31" id="paren.23"/>. As a second region, we analyse temperature variability over the central United States. We hypothesise that these two regions should have different requirements for the ensemble size, with a smaller required ensemble size for the central United States than the tropical Pacific to stay within an acceptable error range.</p>
      <p id="d1e835">For the following examples we use the 2000-year pre-industrial control simulation from the MPI-GE. The advantage of this approach, in contrast to the examples for the forced response, is that the required ensemble size can be estimated for any model without needing a large ensemble to be available. The disadvantage is that, when using the pre-industrial control simulation, we assume that internal variability does not change under global warming.</p>
      <p id="d1e838">We quantify ENSO variability by using the December, January, and February (DJF) variability in the Niño3.4 box (5<inline-formula><mml:math id="M24" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N–5<inline-formula><mml:math id="M25" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> S, 170–120<inline-formula><mml:math id="M26" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> W). To ensure that ENSO variability on interannual to multi-decadal timescales is sampled, we use the Niño3.4 standard deviation for a 100-year period. The standard deviation, as computed for the full 2000-year time series, is used as the truth in this context and indicated by the horizontal black line in Fig. <xref ref-type="fig" rid="Ch1.F7"/>a. To generate synthetic ensemble members, we split the pre-industrial control simulation into overlapping 100-year segments. Each segment is used as one ensemble member, and the temporal standard deviation over the 100-year segment represents ENSO variability for this member. For an ensemble size of one, the spread in ENSO variability seen in Fig. <xref ref-type="fig" rid="Ch1.F7"/>a indicates that individual 100-year periods can have substantially more or less variability than the reference value based on the full control run.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F7" specific-use="star"><?xmltex \currentcnt{7}?><label>Figure 7</label><caption><p id="d1e875">We show for increasing ensemble sizes the <bold>(a)</bold> ENSO variability in the Niño3.4 box (5<inline-formula><mml:math id="M27" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N–5<inline-formula><mml:math id="M28" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> S, 170–120<inline-formula><mml:math id="M29" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> W) calculated over 100-year periods, <bold>(b)</bold> central United States variability (34–46<inline-formula><mml:math id="M30" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> N, 116–96<inline-formula><mml:math id="M31" display="inline"><mml:msup><mml:mi/><mml:mo>∘</mml:mo></mml:msup></mml:math></inline-formula> W) calculated over 100-year periods, <bold>(c)</bold> ENSO variability in the Niño3.4 box calculated over 30-year periods, and <bold>(d)</bold> central United States variability calculated over 30-year periods. All indices are calculated from the 2000-year MPI-GE control run. Each index is calculated as a running value at each time step in the control. ENSO indices are calculated for DJF, and United States indices are calculated for the annual mean. Ensembles of 1 to 120 members are created by randomly sampling the control simulation without replacement. For each ensemble size we create 1000 artificial ensembles. The estimated true value is calculated by using the entire 2000 years of the control and is shown as a horizontal black line. The maximum and minimum values of each index from the 1000 samples are shown as solid black lines. Varying error thresholds are shown as horizontal coloured lines.</p></caption>
          <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://esd.copernicus.org/articles/11/885/2020/esd-11-885-2020-f07.png"/>

        </fig>

      <p id="d1e942">To account for this centennial modulation of ENSO variability, the ENSO variability in multiple ensemble members can be averaged to get a more accurate estimate of the average ENSO variability. We simulate different ensemble sizes by averaging over randomly chosen members for a given ensemble size and repeat this 1000 times. By using a five-member mean, the error of the estimated variability in all samples is within <inline-formula><mml:math id="M32" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">15</mml:mn></mml:mrow></mml:math></inline-formula> % of the true value. To reduce the error below <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula> %, 10 ensemble members are sufficient. To improve the accuracy so that the ENSO variability estimate is within <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> % of the truth, nearly 50 ensemble members are necessary.</p>
      <p id="d1e975">For a region with less variability, much smaller ensemble sizes are sufficient to obtain a similar accuracy. For annual mean central US temperatures (Fig. <xref ref-type="fig" rid="Ch1.F7"/>b) any individual realisation is within <inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">15</mml:mn></mml:mrow></mml:math></inline-formula> % of the truth and 10 members are sufficient to increase the accuracy to the <inline-formula><mml:math id="M36" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> % range around the truth, whereas 50 members are necessary for ENSO. This emphasises that for some regions and quantities a moderate ensemble size or even a single realisation can be sufficient to quantify internal variability.</p>
      <p id="d1e1000">In both examples, the long sampling period of 100 years increases the sample size and thereby improves the accuracy for individual realisations. This is useful if the objective is to quantify variability when stationarity can be assumed, but it can be problematic if the objective is to identify a change in variability, such as changes in ENSO characteristics under global warming. A more detailed discussion of estimating ENSO variability, in particular using the ensemble dimension instead of the time dimension in transient simulations to quantify internal variability, can be found in <xref ref-type="bibr" rid="bib1.bibx20" id="text.24"/> and <xref ref-type="bibr" rid="bib1.bibx16" id="text.25"/>.</p>
<sec id="Ch1.S4.SS2.SSS1">
  <label>4.2.1</label><title>Notes on sampling from a pre-industrial control simulation</title>
      <p id="d1e1016">Sampling from a pre-industrial control simulation to estimate the required ensemble size has two advantages: this can be done before producing a large ensemble for the model and is based on a simulation that is available for every climate model in CMIP5 and CMIP6. Different approaches can be used when sampling from a pre-industrial control simulation. In the following, we discuss different options and their advantages and disadvantages.
<list list-type="bullet"><list-item>
      <p id="d1e1021"><italic>Overlapping segments (applied here)</italic>: we choose to use continuous 100- and 30-year segments to keep temporal autocorrelation intact. From the 2000-year simulation, we can thus generate 20 independent, non-overlapping synthetic realisations (for 100-year segments). To increase the sample size, we allow overlapping segments. These samples are not independent, which leads to a biased estimate, as discussed in Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/>, but enables estimates for ensemble sizes larger than 20.</p></list-item><list-item>
      <p id="d1e1029"><italic>Non-overlapping segments</italic>: the advantage of this approach is that synthetic members can be assumed to be independent and temporal autocorrelation is kept intact. However, for long segments or a short pre-industrial control simulation, only a small number of synthetic members can be generated.</p></list-item><list-item>
      <p id="d1e1035"><italic>Random year selection to generate synthetic segments or members</italic>: the synthetic segments generated by random year selection allow for a wider variety of samples in a segment than continuous segments sampled from<?pagebreak page893?> the pre-industrial control simulation. However, information about temporal autocorrelation is lost, and synthetic segments could have larger variability than continuous segments in the presence of strong  variability on timescales longer than the segment. If the timescale of variability is not the focus of a study, sampling random years to generate synthetic ensemble members can be informative to estimate how well statistics computed across ensemble members <xref ref-type="bibr" rid="bib1.bibx20 bib1.bibx16" id="paren.26"><named-content content-type="pre">e.g.</named-content></xref> capture the model characteristics.</p></list-item></list></p>
</sec>
</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>Quantifying changes in internal variability</title>
      <p id="d1e1054">To quantify changes in internal variability, we need a robust estimate of internal variability both for a reference period and for a period where we want to investigate a potential change in variability (e.g. a pre-industrial control state and a time period in a future scenario). This problem is more challenging than the previous examples because the errors for the variability estimates of the two time periods add up. To demonstrate this, we use the internal variability of September Arctic sea ice area as an example. Previous work has shown that the internal variability in Arctic sea ice area first increases under warming, before it approaches zero when most of the Arctic sea ice has melted <xref ref-type="bibr" rid="bib1.bibx15 bib1.bibx22" id="paren.27"/>. We analyse the 100 members from the 1 % <inline-formula><mml:math id="M37" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">CO</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> scenario from the MPI-GE and use the ensemble standard deviation as an estimator of internal variability. After 120 years, nearly all ensemble members show a completely ice-free Arctic in September (Fig. <xref ref-type="fig" rid="App1.Ch1.S2.F12"/>a). The internal variability increases from model year 1 to year 80, before it sharply drops, reaching zero around year 120 (Fig. <xref ref-type="fig" rid="App1.Ch1.S2.F12"/>b).</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F8" specific-use="star"><?xmltex \currentcnt{8}?><label>Figure 8</label><caption><p id="d1e1077">Change in internal variability of September Arctic sea ice is from the first decade to years 71–80 in a 1 % <inline-formula><mml:math id="M38" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">CO</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> experiment. For different ensemble sizes, we compute the ensemble standard deviation and then average for the first decade and years 71–80 before computing the difference. Black lines show maximum and minimum change in variability from 1000 random samples. Errors are shown as percentage of the 100-member variability change.</p></caption>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://esd.copernicus.org/articles/11/885/2020/esd-11-885-2020-f08.png"/>

        </fig>

      <p id="d1e1097">Here we focus on the increase in variability from the beginning of the simulation to year 80 and ask how many ensemble members are necessary to robustly quantify this change in internal variability. To increase the sample size,<?pagebreak page894?> we use a decadal mean of the ensemble standard deviation rather than a single year. We then compute the difference in internal variability between the two time periods for ensemble sizes between 3 and 100 members. Figure <xref ref-type="fig" rid="Ch1.F8"/> shows the range of this change in internal variability from 1000 random samples. To quantify the change in variability within <inline-formula><mml:math id="M39" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">15</mml:mn></mml:mrow></mml:math></inline-formula> % of the true value (here defined as the internal variability change estimated with 100 members), 50 ensemble members are necessary. An error of less than <inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula> % and <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> % is only reached beyond 50 members. Due to the effect of resampling beyond 50 members, we cannot estimate the required ensemble size for these error thresholds from the 100-member ensemble used here. For very small ensemble sizes, the estimate of the variability change may even show the opposite sign of the true change, i.e. a decrease in internal variability.</p>
      <p id="d1e1133">The large number of ensemble members required to robustly quantify this change in variability shows that identifying a change in internal variability requires the largest ensemble size of all examples shown in this study, even when using decadal averaging to increase the sample size. This is because a robust estimate of a change in internal variability requires a clean separation of internal variability from the forced response and a robust estimate of internal variability for two different time periods. Errors in any of these estimates will propagate to the estimated change in variability, thereby making it more challenging. A small forced change in internal variability will further complicate this analysis.</p>
      <p id="d1e1136">A first estimate for the magnitude of a detectable change in internal variability can be derived from the control run (as in Fig. <xref ref-type="fig" rid="Ch1.F7"/>). Any change in variability that is smaller than the uncertainty of the estimated internal variability for a given ensemble size is not detectable. We note that this method can also be used to add error bars to estimates of forced changes in internal variability under climate change in small ensembles or single realisations from CMIP and hence determine the robustness of results.</p>
</sec>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <label>5</label><title>Summary and conclusions</title>
      <p id="d1e1151">Multiple ensemble members for a single climate model are required for robustly estimating the model's forced response to an external forcing change and its internal variability. Without a robust characterisation of these model characteristics, differences between models or a model and observations can easily be misinterpreted as significant differences, while they could be simply caused by an insufficient sample size. Therefore it is important to use an ensemble size that is sufficiently large to allow a robust quantification of the model characteristic that is investigated.</p>
      <p id="d1e1154">Here we present a generalised approach to estimate the ensemble size that is required to robustly estimate a model's characteristics. While the focus of this study is on the generalised method, the example applications can provide some insight into the required ensemble size for a variety of applications in the MPI-GE. We differentiate three types of question: identifying a forced response, quantifying internal variability, and identifying a change in internal variability. In a next step, an adequate error metric for quantifying the deviations from the true model characteristics is defined, and an acceptable error suitable for the application is chosen. By subsampling a pre-industrial control simulation or a large ensemble of transient simulations, the error for different ensemble sizes can be estimated. By applying the previously selected acceptable error as a threshold to these error estimates for different ensemble sizes, the minimum required ensemble size for the<?pagebreak page895?> given question and model can be determined. Because the subsampling of the full sample does not generate independent samples when approaching the full ensemble size, the error estimate is biased for ensemble sizes close to the available ensemble size. We demonstrate that this resampling effect substantially affects the error estimate when using more than 50 % of the full ensemble. For example, a 50-member ensemble cannot be used to conclude that 50 members are sufficient for a given application, because all ensemble estimates beyond 25 members would be affected by resampling and therefore biased.</p>
      <p id="d1e1157">We apply the method to several examples and use the 200-member historical ensemble, a 2000-year pre-industrial control simulation, and a 100-member 1 % <inline-formula><mml:math id="M42" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">CO</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> experiment from the MPI-GE to estimate required ensemble sizes for various applications for the MPI-ESM model.</p>
      <p id="d1e1171">To identify the externally forced temperature response from 1850 to 2005, most ocean regions require fewer than 10 members, while land regions at higher latitudes may require more than 50 members. To characterise rainfall changes over the same period, more ensemble members are required in the tropics than at higher latitudes. While regions that require more ensemble members can be objectively identified, the required number of members depends on a subjective choice of the acceptable error and can therefore vary substantially for different applications.</p>
      <p id="d1e1175">The analysis of the forced cooling after a volcanic eruption and the analysis of ENSO variability demonstrate that a small ensemble size can lead to a misinterpretation. For the example of the volcanic eruption, an ensemble consisting of two or three members could show a warming after the volcanic eruption, while the true forced response of the model is a cooling. For ENSO, a too-small ensemble still contains a large uncertainty in the estimate of ENSO variability. This may lead to a misinterpretation of a signal as a forced change in ENSO, whereas it might still be within sampling uncertainty.
<xref ref-type="bibr" rid="bib1.bibx31" id="text.28"/> show that samples from different time periods in a pre-industrial control simulation can show substantially different ENSO characteristics. <xref ref-type="bibr" rid="bib1.bibx4" id="text.29"/> on the other hand use single realisations for different models to identify forced changes in ENSO in future projections. While the robustness of the results seems clear given most models show an increase in ENSO amplitude, we show that within a single model differences between realisations can be large due to internal variability alone. By using the method introduced in this study, we can add to the robustness of studies such as <xref ref-type="bibr" rid="bib1.bibx4" id="text.30"/> by adding error bars from the pre-industrial control simulation to each model to test if changes in variability are indeed robust within each model.</p>
      <p id="d1e1187"><?xmltex \hack{\newpage}?>The examples in this study demonstrate that for some applications ensemble sizes around five members are sufficient, while other applications require ensemble sizes well above 100 members. In Sect. <xref ref-type="sec" rid="Ch1.S1"/> we introduced several estimates for required ensemble sizes from the literature. While most of the applications from previous studies are not directly comparable to the examples we use here, the large range of required ensemble sizes emphasises the need to systematically estimate the required ensemble size for each individual application. Furthermore, the required ensemble size may be model dependent. Therefore, the numbers derived in this and previous studies should only be used as approximate estimates and supported by a systematic model- and application-specific estimate following the approach outlined in this study.</p>
      <p id="d1e1193">Information about the sufficient ensemble size is not only crucial when choosing or designing a large ensemble, but it can also help to identify applications where a small number of ensemble members is sufficient and thereby inform the design of multi-model intercomparison studies. The method introduced in this study can add to the robustness of results both from single-model large ensembles and multi-model ensembles.</p><?xmltex \hack{\clearpage}?>
</sec>

      
      </body>
    <back><app-group>

<?pagebreak page896?><app id="App1.Ch1.S1">
  <?xmltex \currentcnt{A}?><label>Appendix A</label><title>Notes on sampling</title>
      <p id="d1e1208">In this study, we made several choices on how we sample from a large ensemble or pre-industrial control simulation. In this section, we discuss alternative sampling approaches and caveats.</p>
<sec id="App1.Ch1.S1.SS1">
  <label>A1</label><title>Resampling with and without replacement</title>
      <p id="d1e1218">We choose to resample without replacement for all examples shown. While this choice leads to ambiguities in error convergence as discussed in Sect. <xref ref-type="sec" rid="App1.Ch1.S1.SS2"/>, we argue that sampling without replacement is a better proxy for what we try to imitate by resampling: a random set of members that we could have produced when running a given number of realisations. Sampling with replacement would mean that for example a randomly sampled five-member ensemble could contain two (or more) identical realisations. Given how SMILEs are initialised, this is unlikely to happen, and, even if it did happen, such an ensemble would not be used as a set of independent realisations without careful investigation.</p>
      <p id="d1e1223">In Fig. <xref ref-type="fig" rid="App1.Ch1.S1.F9"/>, we repeat the analysis shown in Fig. <xref ref-type="fig" rid="Ch1.F2"/> but allow replacement when resampling from the 200 members. We still use the 200-member mean as the reference for the forced response in historical GSAT. Sampling with replacement results in a consistently larger error estimate for the mean RMSE, resulting in a larger required ensemble size for a given error.</p>
</sec>
<sec id="App1.Ch1.S1.SS2">
  <label>A2</label><title>How resampling from a small ensemble can bias the error estimate</title>
      <p id="d1e1238">Generating samples without replacement as applied in this study can bias the error estimate when approaching the full ensemble size. We use the distribution parameters of the full ensemble, for example the mean or standard deviation, as the “truth” in many of the examples shown here. When the size of the sample approaches the size of the full ensemble, for example 190 members from a 200-member ensemble, the difference between these ensembles will be small because they share most of their members. This results in a small error estimate but does not necessarily mean that 190 members are sufficient for a given application.</p>
      <p id="d1e1241">The resampling problem occurs with any limited sample. At some point, the 1000 random subsamples are not independent anymore because they share many of the randomly drawn members from the full ensemble. Therefore, they look more similar not only to each other but also to the 200-member mean. To demonstrate how this resampling affects our estimate of the error, we deliberately reduce the size of the ensemble. For instance, by only using the first 150 members and repeating the analysis (purple line in Fig. <xref ref-type="fig" rid="App1.Ch1.S1.F10"/>), the random samples are subsets of these 150 members. Because the 150-member mean is now used as the best estimate, the RMSE is – by construction – 0 at 150 members. Similar behaviour can be seen when only using the first 100 (red), 75 (green), 50 (blue), and 20 members (yellow line).</p>
      <p id="d1e1246">We investigate at which sample sizes the reduction of the error mainly occurs because of an increased ensemble size, or simply because of resampling that leads to an error convergence without additional information about a sufficient ensemble size. For a smaller number of realisations in the full ensemble, the resampling starts to dominate the error convergence earlier than in a much larger ensemble. Therefore, the comparison of the different maximum ensemble sizes in Fig. <xref ref-type="fig" rid="App1.Ch1.S1.F10"/> indicates when the resampling begins to affect the error convergence. For ensemble sizes that are much smaller than the maximum ensemble size, the different random samples are largely independent and therefore hardly affected by resampling. When increasing the ensemble size in the subsamples, the resampling starts to affect the error estimate for a small maximum ensemble size (e.g. 20 members), whereas the samples are still independent when drawn from a much larger maximum ensemble size (e.g. 200 members). The sample size for which the RMSE estimate in a smaller maximum ensemble size starts to diverge from the RMSE estimate based on a larger maximum ensemble size determines the threshold of where resampling substantially affects the error convergence. Beyond this sample size, the error estimate should not be used to approximate the true error.</p>
      <p id="d1e1251">We find that the RMSE estimates for different maximum ensemble sizes in Fig. <xref ref-type="fig" rid="App1.Ch1.S1.F10"/> always start to diverge when about 50% of the maximum ensemble size is used. This implies that up to 50 % of the maximum ensemble size can be used to estimate the forced response of GSAT in a transient forcing scenario without a major impact from resampling.</p>
      <p id="d1e1257">The same resampling problem also occurs for other questions. To demonstrate this, we investigate how many members are necessary to sample ENSO variability. We use the 50-year standard deviation of the Niño3.4 box to quantify ENSO variability. A single 50-year period is treated as one ensemble member. Random subsamples of 50-year periods from the 2000-year pre-industrial control simulation from the MPI-GE are used to generate a synthetic ensemble. In Fig. <xref ref-type="fig" rid="App1.Ch1.S1.F11"/>, the light blue envelope shows that, by averaging the standard deviation from more members, a more accurate estimate of ENSO variability can be obtained.</p>
      <p id="d1e1262">We then reduce the maximum ensemble size by using only 500 (200, 100, and 50) years from the control run. Similar to the result in Fig. <xref ref-type="fig" rid="App1.Ch1.S1.F10"/>, the error appears to converge when approaching the maximum ensemble size. By comparing the different maximum ensemble sizes in Fig. <xref ref-type="fig" rid="App1.Ch1.S1.F11"/>, we can see that the resampling begins to affect the error estimate when the ensemble size approaches 50 % of the maximum ensemble size.</p>
      <p id="d1e1269">These two independent lines of evidence demonstrate that resampling affects the error estimate when using more than 50 % of the available maximum sample size (either ensemble members or years in a pre-industrial control simulation).<?pagebreak page897?> Beyond this ensemble size, the analysis does not provide a realistic estimate of the error and conclusions about the required ensemble size will be biased low. We note that for very simple applications, such as the mean of a stationary time series, the error scales with <inline-formula><mml:math id="M43" display="inline"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:msqrt><mml:mi>n</mml:mi></mml:msqrt></mml:mfrac></mml:mstyle></mml:math></inline-formula>. For more complex error estimates, such as the RMSE between non-stationary time series, the scaling law is not as simple, which is why we rely on the empirical analysis outlined above.</p>

      <?xmltex \floatpos{b!}?><fig id="App1.Ch1.S1.F9"><?xmltex \currentcnt{A1}?><label>Figure A1</label><caption><p id="d1e1286">Sampling with or without replacement affects the error estimate and therefore the estimate for the required ensemble size. The black line shows the mean RMSE for GSAT for ensemble sizes from 2 to 200. The reference is the 200-member mean from Fig. <xref ref-type="fig" rid="Ch1.F1"/>, and the RMSE is computed for all 1000 samples. The shaded area shows the range of RMSE values for individual samples; the solid line shows the mean RMSE. The red line and shading show the RMSE for ensemble sizes from 2 to 200, but samples are generated by allowing sampling with replacement.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://esd.copernicus.org/articles/11/885/2020/esd-11-885-2020-f09.png"/>

        </fig>

      <?xmltex \floatpos{b!}?><fig id="App1.Ch1.S1.F10"><?xmltex \currentcnt{A2}?><label>Figure A2</label><caption><p id="d1e1301">In a smaller ensemble, the RMSE converges to zero earlier. This is caused by resampling and does not indicate that the error is small. The black line shows the mean RMSE for GSAT for ensemble sizes from 2 to 200. The reference is the 200-member mean from Fig. <xref ref-type="fig" rid="Ch1.F1"/>, and the RMSE is computed for all 1000 samples. The shaded area shows the range of RMSE values for individual samples; the solid line shows the mean RMSE. The other colours show the same analysis after excluding the last 50 members (purple), 100 members (red), 125 members (green), 150 members (blue), and 180 members (yellow) from the ensemble.</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=341.433071pt}?><graphic xlink:href="https://esd.copernicus.org/articles/11/885/2020/esd-11-885-2020-f10.png"/>

        </fig>

<?xmltex \hack{\clearpage}?><?xmltex \floatpos{h!}?><fig id="App1.Ch1.S1.F11"><?xmltex \currentcnt{A3}?><label>Figure A3</label><caption><p id="d1e1318">Probability density function (PDF) of ensemble-averaged Niño3.4 standard deviations possible in the MPI-GE pre-industrial control simulation for subsampling ensembles ranging from 50 to 1000 members (shown as different colours) for smaller ensemble sizes. Each PDF is shown relative to the corresponding ensemble mean value. We use the last 1000 years of the 2000-year control run to calculate the ranges. The Niño3.4 standard deviation is calculated over 50-year periods. The PDFs are created by resampling the control simulation 1000 times. For each PDF the entirety of the 1000 years is used (i.e. the blue 500-member PDF is the mean of two 500-member PDFs).</p></caption>
          <?xmltex \hack{\hsize\textwidth}?>
          <?xmltex \igopts{width=236.157874pt}?><graphic xlink:href="https://esd.copernicus.org/articles/11/885/2020/esd-11-885-2020-f11.png"/>

        </fig>

<?xmltex \hack{\clearpage}?>
</sec>
</app>

<?pagebreak page899?><app id="App1.Ch1.S2">
  <?xmltex \currentcnt{B}?><label>Appendix B</label><title>Arctic sea ice area under strong warming</title>
      <p id="d1e1340">The internal variability of September Arctic sea ice area is known to change under global warming. In this study, we use September Arctic sea ice area as an example for a quantity with a change in internal variability under global warming.</p>
      <p id="d1e1343">Previous work has shown that the internal variability in Arctic sea ice area first increases under warming, before it approaches zero when most of the Arctic sea ice has melted <xref ref-type="bibr" rid="bib1.bibx15 bib1.bibx22" id="paren.31"/>. We analyse the 100 members from the 1 % <inline-formula><mml:math id="M44" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">CO</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> scenario from the MPI-GE and use the ensemble standard deviation as an estimator of internal variability. After 120 years, nearly all ensemble members show a completely ice-free Arctic in September (Fig. <xref ref-type="fig" rid="App1.Ch1.S2.F12"/>a). The internal variability increases from model year 1 to year 80, before it sharply drops, reaching zero around year 120 when all sea ice is lost (Fig. <xref ref-type="fig" rid="App1.Ch1.S2.F12"/>b).</p>

      <?xmltex \floatpos{b!}?><fig id="App1.Ch1.S2.F12"><?xmltex \currentcnt{B1}?><label>Figure B1</label><caption><p id="d1e1366"><bold>(a)</bold> September Arctic sea ice area in the 100 realisations for the 1 % <inline-formula><mml:math id="M45" display="inline"><mml:mrow class="chem"><mml:msub><mml:mi mathvariant="normal">CO</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> experiment. <bold>(b)</bold> Ensemble standard deviation for the 100 realisations.</p></caption>
        <?xmltex \hack{\hsize\textwidth}?>
        <?xmltex \igopts{width=312.980315pt}?><graphic xlink:href="https://esd.copernicus.org/articles/11/885/2020/esd-11-885-2020-f12.png"/>

      </fig>

<?xmltex \hack{\clearpage}?>
</app>
  </app-group><notes notes-type="codeavailability"><title>Code availability</title>

      <p id="d1e1399">Primary data and scripts used in the analysis and other supporting information that may be useful in reproducing the author's work are archived by the Max Planck Institute for Meteorology and can be obtained by contacting publications@mpimet.mpg.de.</p>
  </notes><notes notes-type="dataavailability"><title>Data availability</title>

      <p id="d1e1405">Output from the MPI Grand Ensemble that was used in this study and additional output can be downloaded from <uri>https://www.mpimet.mpg.de/en/grand-ensemble/</uri> (last access: October 2020). The output is described in <ext-link xlink:href="https://doi.org/10.1029/2019MS001639" ext-link-type="DOI">10.1029/2019MS001639</ext-link> <xref ref-type="bibr" rid="bib1.bibx21" id="paren.32"/>.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d1e1420">All authors conceptualised the study and carried out the formal analysis. SM wrote the original draft with input from all authors.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d1e1426">The authors declare that they have no conflict of interest.</p>
  </notes><notes notes-type="sistatement"><title>Special issue statement</title>

      <p id="d1e1432">This article is part of the special issue “Large Ensemble Climate Model Simulations: Exploring Natural Variability, Change Signals and Impacts”. It is not associated with a conference.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e1438">We thank Chao Li for conducting an internal review of the manuscript, Jin-Song von Storch for helpful comments, and the Max Planck Society for the Advancement of Science for funding all three authors. We thank the two anonymous reviewers for constructive feedback. We thank Mikhail Dobrynin and Johanna Baehr from the University of Hamburg for completing the second 100 MPI-GE ensemble simulations and providing the data from these simulations for use in this paper. Dirk Olonscheck was supported by the European Union's Horizon 2020 Research and Innovation Programme under grant agreement number 820829 (CONSTRAIN).</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d1e1443">This research has been supported by the European Union Horizon 2020 Research and Innovation Programme (grant no. 820829 (CONSTRAIN)). <?xmltex \hack{\newline}?><?xmltex \hack{\newline}?> The article processing charges for this open-access <?xmltex \hack{\newline}?> publication were covered by the Max Planck Society.</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d1e1454">This paper was edited by Ralf Ludwig and reviewed by two anonymous referees.</p>
  </notes><?xmltex \hack{\newpage}?><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Bellenger et al.(2013)Bellenger, Guilyardi, Leloup, Lengaigne, and
Vialard</label><?label Bellenger:2013jr?><mixed-citation>
Bellenger, H., Guilyardi, É., Leloup, J., Lengaigne, M., and Vialard, J.:
ENSO representation in climate models: from CMIP3 to CMIP5, Clim. Dynam., 42, 1999–2018, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Bittner et al.(2016)Bittner, Schmidt, Timmreck, and
Sienz</label><?label Bittner:2016kk?><mixed-citation>
Bittner, M., Schmidt, H., Timmreck, C., and Sienz, F.: Using a large ensemble of simulations to assess the Northern Hemisphere stratospheric dynamical
response to tropical volcanic eruptions and its uncertainty, Geophys. Res. Lett., 43, 9324–9332, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Branstator and Selten(2009)</label><?label branstator_selten2009?><mixed-citation>Branstator, G. and Selten, F.: `Modes of Variability' and Climate Change, J. Climate, 22, 2639–2658, <ext-link xlink:href="https://doi.org/10.1175/2008JCLI2517.1" ext-link-type="DOI">10.1175/2008JCLI2517.1</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx4"><?xmltex \def\ref@label{{Cai et~al.(2018)Cai, Wang, Dewitte, Wu, Santoso, Takahashi, Yang,
Carr{\'{e}}ric, and McPhaden}}?><label>Cai et al.(2018)Cai, Wang, Dewitte, Wu, Santoso, Takahashi, Yang,
Carréric, and McPhaden</label><?label Cai:2018gn?><mixed-citation>
Cai, W., Wang, G., Dewitte, B., Wu, L., Santoso, A., Takahashi, K., Yang, Y.,
Carréric, A., and McPhaden, M. J.: Increased variability of eastern
Pacific El Niño under greenhouse warming, Nature, 564, 1–18, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Christensen et al.(2013)Christensen, H., Aldrian, An, Cavalcanti,
de Castro, Dong, Goswami, Hall, Kanyanga, Kitoh, Kossin, Lau, Renwick,
Stephenson, Xie, and Zhou</label><?label christensen_etal2013?><mixed-citation>
Christensen, J. H. K. K. K., Aldrian, E., An, S.-I., Cavalcanti, I.,
de Castro, M., Dong, W., Goswami, P., Hall, A., Kanyanga, J., Kitoh, A.,
Kossin, J., Lau, N.-C., Renwick, J., Stephenson, D., Xie, S.-P., and Zhou,
T.: Climate Phenomena and their Relevance for Future Regional Climate Change, in: Climate Change 2013: The Physical Science Basis, Contribution of Working Group I to the Fifth Assessment Report of the Intergovernmental Panel on Climate Change, edited by: Stocker, T. F., Qin, D., Plattner, G.-K., Tignor, M., Allen, S., Boschung, J., Nauels, A., Xia, Y., Bex, V., and Midgley, P., Cambridge University Press, Cambridge, 1217–1308, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Daron and Stainforth(2013)</label><?label Daron_2013?><mixed-citation>Daron, J. D. and Stainforth, D. A.: On predicting climate under climate change, Environ. Res. Lett., 8, 034021, <ext-link xlink:href="https://doi.org/10.1088/1748-9326/8/3/034021" ext-link-type="DOI">10.1088/1748-9326/8/3/034021</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Deser et al.(2012a)Deser, Knutti, Solomon, and
Phillips</label><?label Deser:2012iy?><mixed-citation>
Deser, C., Knutti, R., Solomon, S., and Phillips, A. S.: Communication of the
role of natural variability in future North American climate, Nat. Climate
Change, 2, 775–779, 2012a.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Deser et al.(2012b)Deser, Phillips, Bourdette, and
Teng</label><?label Deser:2012fa?><mixed-citation>
Deser, C., Phillips, A., Bourdette, V., and Teng, H.: Uncertainty in climate
change projections: the role of internal variability, Clim. Dynam., 38, 527–546, 2012b.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Deser et al.(2020)Deser, Lehner, Rodgers, Ault, Delworth, diNezio,
Fiore, Frankignoul, Fyfe, Horton, Kay, Knutti, Lovenduski, Marotzke,
McKinnon, Minobe, Randerson, Screen, Simpson, and Ting</label><?label Deser:2019?><mixed-citation>Deser, C., Lehner, F., Rodgers, K. B., Ault, T. R., Delworth, T. L., diNezio,
P., Fiore, A., Frankignoul, C., Fyfe, J. C., Horton, D. E., Kay, J. E., Knutti, R., Lovenduski, N. S., Marotzke, J., McKinnon, K. A., Minobe, S.,
Randerson, J., Screen, J. A., Simpson, I. R., and Ting, M.: Strength in
Numbers: The Utility of Large Ensembles with Multiple Earth System Models,
Nat. Clim. Change, 10, 277–286, <ext-link xlink:href="https://doi.org/10.1038/s41558-020-0731-2" ext-link-type="DOI">10.1038/s41558-020-0731-2</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx10"><?xmltex \def\ref@label{{Dr\'{o}tos et~al.(2017)Dr{\'{o}}tos, B{\'{o}}dai, and
T{\'{e}}l}}?><label>Drótos et al.(2017)Drótos, Bódai, and
Tél</label><?label drotos_etal2017?><mixed-citation>Drótos, G., Bódai, T., and Tél, T.: On the importance of the
convergence to climate attractors, Eur. Phys. J. Spec. Top., 226, 2031–2038, <ext-link xlink:href="https://doi.org/10.1140/epjst/e2017-70045-7" ext-link-type="DOI">10.1140/epjst/e2017-70045-7</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Ferro et al.(2012)Ferro, Jupp, Lambert, Huntingford, and
Cox</label><?label Ferro:2012hz?><mixed-citation>
Ferro, C. A. T., Jupp, T. E., Lambert, F. H., Huntingford, C., and Cox, P. M.: Model complexity versus ensemble size: allocating resources for climate
prediction, Philos. T. Roy. Soc. A, 370, 1087–1099, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Frankcombe et al.(2015)Frankcombe, England, Mann, and
Steinman</label><?label Frankcombe:2015kx?><mixed-citation>
Frankcombe, L. M., England, M. H., Mann, M. E., and Steinman, B. A.: Separating Internal Variability from the Externally Forced Climate Response, J. Climate, 28, 8184–8202, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Frankcombe et al.(2018)Frankcombe, England, Kajtar, Mann, and
Steinman</label><?label Frankcombe:2018bb?><mixed-citation>
Frankcombe, L. M., England, M. H., Kajtar, J. B., Mann, M. E., and Steinman,
B. A.: On the Choice of Ensemble Mean for Estimating the Forced Signal in
the Presence of Internal Variability, J. Climate, 31, 5681–5693, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Frankignoul et al.(2017)Frankignoul, Gastineau, and
Kwon</label><?label frankignoul_etal2017?><mixed-citation>Frankignoul, C., Gastineau, G., and Kwon, Y.-O.:  Estimation of the SST
Response to Anthropogenic <?pagebreak page901?>and External Forcing and Its Impact on the Atlantic
Multidecadal Oscillation and the Pacific Decadal Oscillation, J. Climate, 30, 9871–9895, <ext-link xlink:href="https://doi.org/10.1175/JCLI-D-17-0009.1" ext-link-type="DOI">10.1175/JCLI-D-17-0009.1</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Goosse et al.(2009)Goosse, Arzel, Bitz, de Montety, and
Vancoppenolle</label><?label Goosse:2009dz?><mixed-citation>Goosse, H., Arzel, O., Bitz, C. M., de Montety, A., and Vancoppenolle, M.:
Increased variability of the Arctic summer ice extent in a warmer climate,
Geophys. Res. Lett., 36, L23702, <ext-link xlink:href="https://doi.org/10.1029/2009GL040546" ext-link-type="DOI">10.1029/2009GL040546</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx16"><?xmltex \def\ref@label{{Haszpra et~al.(2020)Haszpra, Herein, and B{\'{o}}dai}}?><label>Haszpra et al.(2020)Haszpra, Herein, and Bódai</label><?label Haszpra:2020ce?><mixed-citation>Haszpra, T., Herein, M., and Bódai, T.: Investigating ENSO and its teleconnections under climate change in an ensemble view – a new perspective, Earth Syst. Dynam., 11, 267–280, <ext-link xlink:href="https://doi.org/10.5194/esd-11-267-2020" ext-link-type="DOI">10.5194/esd-11-267-2020</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Kay et al.(2015)Kay, Deser, Phillips, Mai, Hannay, Strand, Arblaster, Bates, Danabasoglu, Edwards, Holland, Kushner, Lamarque, Lawrence, Lindsay, Middleton, Munoz, Neale, Oleson, Polvani, and Vertenstein</label><?label Kay:2015jh?><mixed-citation>
Kay, J. E., Deser, C., Phillips, A., Mai, A., Hannay, C., Strand, G.,
Arblaster, J. M., Bates, S. C., Danabasoglu, G., Edwards, J., Holland, M.,
Kushner, P., Lamarque, J. F., Lawrence, D., Lindsay, K., Middleton, A.,
Munoz, E., Neale, R., Oleson, K., Polvani, L., and Vertenstein, M.: The
Community Earth System Model (CESM) Large Ensemble Project: A Community
Resource for Studying Climate Change in the Presence of Internal Climate
Variability, B. Am. Meteorol. Soc., 96, 1333–1349, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Kirchmeier-Young et al.(2017)Kirchmeier-Young, Zwiers, and
Gillett</label><?label kirchmeier_etal2017?><mixed-citation>Kirchmeier-Young, M. C., Zwiers, F. W., and Gillett, N. P.: Attribution of
Extreme Events in Arctic Sea Ice Extent, J. Climate, 30, 553–571,
<ext-link xlink:href="https://doi.org/10.1175/JCLI-D-16-0412.1" ext-link-type="DOI">10.1175/JCLI-D-16-0412.1</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Li and Ilyina(2018)</label><?label Li:2018jx?><mixed-citation>
Li, H. and Ilyina, T.: Current and Future Decadal Trends in the Oceanic Carbon Uptake Are Dominated by Internal Variability, Geophys. Res. Lett., 45, 916–925, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Maher et al.(2018)Maher, Matei, Milinski, and
Marotzke</label><?label Maher:2018iy?><mixed-citation>
Maher, N., Matei, D., Milinski, S., and Marotzke, J.: ENSO change in climate
projections: forced response or internal variability?, Geophys. Res. Lett., 45, 11390–11398, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx21"><?xmltex \def\ref@label{{Maher et~al.(2019)Maher, Milinski, Su{\'{a}}rez-Guti{\'{e}}rrez, Botzet, Dobrynin, Kornblueh, Kr{\"{o}}ger, Takano, Ghosh, Hedemann, Li, Li, Manzini, Notz, Putrasahan, Boysen, Claussen, Ilyina, Olonscheck, Raddatz, Stevens, and Marotzke}}?><label>Maher et al.(2019)Maher, Milinski, Suárez-Gutiérrez, Botzet, Dobrynin, Kornblueh, Kröger, Takano, Ghosh, Hedemann, Li, Li, Manzini, Notz, Putrasahan, Boysen, Claussen, Ilyina, Olonscheck, Raddatz, Stevens, and Marotzke</label><?label Maher:2019dw?><mixed-citation>Maher, N., Milinski, S., Suárez-Gutiérrez, L., Botzet, M., Dobrynin,
M., Kornblueh, L., Kröger, J., Takano, Y., Ghosh, R., Hedemann, C., Li,
C., Li, H., Manzini, E., Notz, D., Putrasahan, D., Boysen, L., Claussen, M.,
Ilyina, T., Olonscheck, D., Raddatz, T., Stevens, B., and Marotzke, J.: The
Max Planck Institute Grand Ensemble: Enabling the Exploration of Climate
System Variability, J. Adv. Model. Earth Syst., 28, 867–920, <ext-link xlink:href="https://doi.org/10.1029/2019MS001639" ext-link-type="DOI">10.1029/2019MS001639</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Olonscheck and Notz(2017)</label><?label olonscheck_notz2017?><mixed-citation>Olonscheck, D. and Notz, D.: Consistently Estimating Internal Climate
Variability from Climate Model Simulations, J. Climate, 30, 9555–9573, <ext-link xlink:href="https://doi.org/10.1175/JCLI-D-16-0428.1" ext-link-type="DOI">10.1175/JCLI-D-16-0428.1</ext-link>, 2017.
</mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx23"><label>Pausata et al.(2015)Pausata, Grini, Caballero, Hannachi, and
Seland</label><?label Pausata:2015hz?><mixed-citation>
Pausata, F. S. R., Grini, A., Caballero, R., Hannachi, A., and Seland, Ø.:
High-latitude volcanic eruptions in the Norwegian Earth System Model: the
effect of different initial conditions and of the ensemble size, Tellus B, 11, 2050–2069, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx24"><?xmltex \def\ref@label{{Rodgers et~al.(2015)Rodgers, Lin, and Fr{\"{o}}licher}}?><label>Rodgers et al.(2015)Rodgers, Lin, and Frölicher</label><?label Rodgers:2015jq?><mixed-citation>Rodgers, K. B., Lin, J., and Frölicher, T. L.: Emergence of multiple ocean ecosystem drivers in a large ensemble suite with an Earth system model, Biogeosciences, 12, 3301–3320, <ext-link xlink:href="https://doi.org/10.5194/bg-12-3301-2015" ext-link-type="DOI">10.5194/bg-12-3301-2015</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Steinman et al.(2015)Steinman, Frankcombe, Mann, Miller, and
England</label><?label Steinman:2015if?><mixed-citation>Steinman, B. A., Frankcombe, L. M., Mann, M. E., Miller, S. K., and England,
M. H.: Response to Comment on “Atlantic and Pacific multidecadal oscillations and Northern Hemisphere temperatures”, Science, 350, 1326, <ext-link xlink:href="https://doi.org/10.1126/science.aac5208" ext-link-type="DOI">10.1126/science.aac5208</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Stevenson et al.(2012)Stevenson, Fox-Kemper, Jochum, Neale, Deser,
and Meehl</label><?label Stevenson:2012hk?><mixed-citation>
Stevenson, S., Fox-Kemper, B., Jochum, M., Neale, R., Deser, C., and Meehl, G.: Will There Be a Significant Change to El Niño in the Twenty-First
Century?, J. Climate, 25, 2129–2145, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx27"><?xmltex \def\ref@label{{Stolpe et~al.(2018)Stolpe, Medhaug, Sedl\'{a}\v{c}ek, and
Knutti}}?><label>Stolpe et al.(2018)Stolpe, Medhaug, Sedláček, and
Knutti</label><?label stolpe_etal2018?><mixed-citation>Stolpe, M. B., Medhaug, I., Sedláček, J., and Knutti, R.: Multidecadal Variability in Global Surface Temperatures Related to the Atlantic Meridional Overturning Circulation, J. Climate, 31, 2889–2906,
<ext-link xlink:href="https://doi.org/10.1175/JCLI-D-17-0444.1" ext-link-type="DOI">10.1175/JCLI-D-17-0444.1</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Taylor et al.(2012)Taylor, Stouffer, and Meehl</label><?label Taylor:2012jg?><mixed-citation>
Taylor, K. E., Stouffer, R. J., and Meehl, G. A.: An Overview of CMIP5 and the Experiment Design, B. Am. Meteorol. Soc., 93, 485–498, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>Thompson et al.(2015)Thompson, Barnes, Deser, Foust, and
Phillips</label><?label Thompson:2015fw?><mixed-citation>
Thompson, D. W. J., Barnes, E. A., Deser, C., Foust, W. E., and Phillips, A. S.: Quantifying the Role of Internal Climate Variability in Future Climate Trends, J. Climate, 28, 6443–6456, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx30"><?xmltex \def\ref@label{{von~K\"{a}nel et~al.(2017){von K\"{a}nel}, Fr\"{o}licher, and
Gruber}}?><label>von Känel et al.(2017)von Känel, Frölicher, and
Gruber</label><?label vonkanel_etal2017?><mixed-citation>von Känel, L., Frölicher, T. L., and Gruber, N.: Hiatus-like decades
in the absence of equatorial Pacific cooling and accelerated global ocean
heat uptake, Geophys. Res. Lett., 44, 7909–7918, <ext-link xlink:href="https://doi.org/10.1002/2017GL073578" ext-link-type="DOI">10.1002/2017GL073578</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Wittenberg(2009)</label><?label Wittenberg:2009ja?><mixed-citation>
Wittenberg, A. T.: Are historical records sufficient to constrain ENSO
simulations?, Geophys. Res. Lett., 36, 3–5, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Zelle et al.(2005)Zelle, Jan van Oldenborgh, Burgers, and
Dijkstra</label><?label zelle_etal2005?><mixed-citation>Zelle, H., van Oldenborgh, G. J., Burgers, G., and Dijkstra, H.: El Niño
and Greenhouse Warming: Results from Ensemble Simulations with the NCAR CCSM, J. Climate, 18, 4669–4683, <ext-link xlink:href="https://doi.org/10.1175/JCLI3574.1" ext-link-type="DOI">10.1175/JCLI3574.1</ext-link>, 2005.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>How large does a large ensemble need to be?</article-title-html>
<abstract-html><p>Initial-condition large ensembles with ensemble sizes ranging from 30 to 100 members have become a commonly used tool for quantifying the forced response and internal variability in various components of the climate system. However, there is no consensus on the ideal or even sufficient ensemble size for a large ensemble. Here, we introduce an objective method to estimate the required ensemble size that can be applied to any given application and demonstrate its use on the examples of global mean near-surface air temperature, local temperature and precipitation, and variability in the El Niño–Southern Oscillation (ENSO) region and central United States for the Max Planck Institute Grand Ensemble (MPI-GE). Estimating the required ensemble size is relevant not only for designing or choosing a large ensemble but also for designing targeted sensitivity experiments with a model. Where possible, we base our estimate of the required ensemble size on the pre-industrial control simulation, which is available for every model. We show that more ensemble members are needed to quantify variability than the forced response, with the largest ensemble sizes needed to detect changes in internal variability itself. Finally, we highlight that the required ensemble size depends on both the acceptable error to the user and the studied quantity.</p></abstract-html>
<ref-html id="bib1.bib1"><label>Bellenger et al.(2013)Bellenger, Guilyardi, Leloup, Lengaigne, and
Vialard</label><mixed-citation>
Bellenger, H., Guilyardi, É., Leloup, J., Lengaigne, M., and Vialard, J.:
ENSO representation in climate models: from CMIP3 to CMIP5, Clim. Dynam., 42, 1999–2018, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Bittner et al.(2016)Bittner, Schmidt, Timmreck, and
Sienz</label><mixed-citation>
Bittner, M., Schmidt, H., Timmreck, C., and Sienz, F.: Using a large ensemble of simulations to assess the Northern Hemisphere stratospheric dynamical
response to tropical volcanic eruptions and its uncertainty, Geophys. Res. Lett., 43, 9324–9332, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Branstator and Selten(2009)</label><mixed-citation>
Branstator, G. and Selten, F.: `Modes of Variability' and Climate Change, J. Climate, 22, 2639–2658, <a href="https://doi.org/10.1175/2008JCLI2517.1" target="_blank">https://doi.org/10.1175/2008JCLI2517.1</a>, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Cai et al.(2018)Cai, Wang, Dewitte, Wu, Santoso, Takahashi, Yang,
Carréric, and McPhaden</label><mixed-citation>
Cai, W., Wang, G., Dewitte, B., Wu, L., Santoso, A., Takahashi, K., Yang, Y.,
Carréric, A., and McPhaden, M. J.: Increased variability of eastern
Pacific El Niño under greenhouse warming, Nature, 564, 1–18, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Christensen et al.(2013)Christensen, H., Aldrian, An, Cavalcanti,
de Castro, Dong, Goswami, Hall, Kanyanga, Kitoh, Kossin, Lau, Renwick,
Stephenson, Xie, and Zhou</label><mixed-citation>
Christensen, J. H. K. K. K., Aldrian, E., An, S.-I., Cavalcanti, I.,
de Castro, M., Dong, W., Goswami, P., Hall, A., Kanyanga, J., Kitoh, A.,
Kossin, J., Lau, N.-C., Renwick, J., Stephenson, D., Xie, S.-P., and Zhou,
T.: Climate Phenomena and their Relevance for Future Regional Climate Change, in: Climate Change 2013: The Physical Science Basis, Contribution of Working Group I to the Fifth Assessment Report of the Intergovernmental Panel on Climate Change, edited by: Stocker, T. F., Qin, D., Plattner, G.-K., Tignor, M., Allen, S., Boschung, J., Nauels, A., Xia, Y., Bex, V., and Midgley, P., Cambridge University Press, Cambridge, 1217–1308, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Daron and Stainforth(2013)</label><mixed-citation>
Daron, J. D. and Stainforth, D. A.: On predicting climate under climate change, Environ. Res. Lett., 8, 034021, <a href="https://doi.org/10.1088/1748-9326/8/3/034021" target="_blank">https://doi.org/10.1088/1748-9326/8/3/034021</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Deser et al.(2012a)Deser, Knutti, Solomon, and
Phillips</label><mixed-citation>
Deser, C., Knutti, R., Solomon, S., and Phillips, A. S.: Communication of the
role of natural variability in future North American climate, Nat. Climate
Change, 2, 775–779, 2012a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Deser et al.(2012b)Deser, Phillips, Bourdette, and
Teng</label><mixed-citation>
Deser, C., Phillips, A., Bourdette, V., and Teng, H.: Uncertainty in climate
change projections: the role of internal variability, Clim. Dynam., 38, 527–546, 2012b.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Deser et al.(2020)Deser, Lehner, Rodgers, Ault, Delworth, diNezio,
Fiore, Frankignoul, Fyfe, Horton, Kay, Knutti, Lovenduski, Marotzke,
McKinnon, Minobe, Randerson, Screen, Simpson, and Ting</label><mixed-citation>
Deser, C., Lehner, F., Rodgers, K. B., Ault, T. R., Delworth, T. L., diNezio,
P., Fiore, A., Frankignoul, C., Fyfe, J. C., Horton, D. E., Kay, J. E., Knutti, R., Lovenduski, N. S., Marotzke, J., McKinnon, K. A., Minobe, S.,
Randerson, J., Screen, J. A., Simpson, I. R., and Ting, M.: Strength in
Numbers: The Utility of Large Ensembles with Multiple Earth System Models,
Nat. Clim. Change, 10, 277–286, <a href="https://doi.org/10.1038/s41558-020-0731-2" target="_blank">https://doi.org/10.1038/s41558-020-0731-2</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Drótos et al.(2017)Drótos, Bódai, and
Tél</label><mixed-citation>
Drótos, G., Bódai, T., and Tél, T.: On the importance of the
convergence to climate attractors, Eur. Phys. J. Spec. Top., 226, 2031–2038, <a href="https://doi.org/10.1140/epjst/e2017-70045-7" target="_blank">https://doi.org/10.1140/epjst/e2017-70045-7</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Ferro et al.(2012)Ferro, Jupp, Lambert, Huntingford, and
Cox</label><mixed-citation>
Ferro, C. A. T., Jupp, T. E., Lambert, F. H., Huntingford, C., and Cox, P. M.: Model complexity versus ensemble size: allocating resources for climate
prediction, Philos. T. Roy. Soc. A, 370, 1087–1099, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Frankcombe et al.(2015)Frankcombe, England, Mann, and
Steinman</label><mixed-citation>
Frankcombe, L. M., England, M. H., Mann, M. E., and Steinman, B. A.: Separating Internal Variability from the Externally Forced Climate Response, J. Climate, 28, 8184–8202, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Frankcombe et al.(2018)Frankcombe, England, Kajtar, Mann, and
Steinman</label><mixed-citation>
Frankcombe, L. M., England, M. H., Kajtar, J. B., Mann, M. E., and Steinman,
B. A.: On the Choice of Ensemble Mean for Estimating the Forced Signal in
the Presence of Internal Variability, J. Climate, 31, 5681–5693, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Frankignoul et al.(2017)Frankignoul, Gastineau, and
Kwon</label><mixed-citation>
Frankignoul, C., Gastineau, G., and Kwon, Y.-O.:  Estimation of the SST
Response to Anthropogenic and External Forcing and Its Impact on the Atlantic
Multidecadal Oscillation and the Pacific Decadal Oscillation, J. Climate, 30, 9871–9895, <a href="https://doi.org/10.1175/JCLI-D-17-0009.1" target="_blank">https://doi.org/10.1175/JCLI-D-17-0009.1</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Goosse et al.(2009)Goosse, Arzel, Bitz, de Montety, and
Vancoppenolle</label><mixed-citation>
Goosse, H., Arzel, O., Bitz, C. M., de Montety, A., and Vancoppenolle, M.:
Increased variability of the Arctic summer ice extent in a warmer climate,
Geophys. Res. Lett., 36, L23702, <a href="https://doi.org/10.1029/2009GL040546" target="_blank">https://doi.org/10.1029/2009GL040546</a>, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Haszpra et al.(2020)Haszpra, Herein, and Bódai</label><mixed-citation>
Haszpra, T., Herein, M., and Bódai, T.: Investigating ENSO and its teleconnections under climate change in an ensemble view – a new perspective, Earth Syst. Dynam., 11, 267–280, <a href="https://doi.org/10.5194/esd-11-267-2020" target="_blank">https://doi.org/10.5194/esd-11-267-2020</a>, 2020.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Kay et al.(2015)Kay, Deser, Phillips, Mai, Hannay, Strand, Arblaster, Bates, Danabasoglu, Edwards, Holland, Kushner, Lamarque, Lawrence, Lindsay, Middleton, Munoz, Neale, Oleson, Polvani, and Vertenstein</label><mixed-citation>
Kay, J. E., Deser, C., Phillips, A., Mai, A., Hannay, C., Strand, G.,
Arblaster, J. M., Bates, S. C., Danabasoglu, G., Edwards, J., Holland, M.,
Kushner, P., Lamarque, J. F., Lawrence, D., Lindsay, K., Middleton, A.,
Munoz, E., Neale, R., Oleson, K., Polvani, L., and Vertenstein, M.: The
Community Earth System Model (CESM) Large Ensemble Project: A Community
Resource for Studying Climate Change in the Presence of Internal Climate
Variability, B. Am. Meteorol. Soc., 96, 1333–1349, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Kirchmeier-Young et al.(2017)Kirchmeier-Young, Zwiers, and
Gillett</label><mixed-citation>
Kirchmeier-Young, M. C., Zwiers, F. W., and Gillett, N. P.: Attribution of
Extreme Events in Arctic Sea Ice Extent, J. Climate, 30, 553–571,
<a href="https://doi.org/10.1175/JCLI-D-16-0412.1" target="_blank">https://doi.org/10.1175/JCLI-D-16-0412.1</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Li and Ilyina(2018)</label><mixed-citation>
Li, H. and Ilyina, T.: Current and Future Decadal Trends in the Oceanic Carbon Uptake Are Dominated by Internal Variability, Geophys. Res. Lett., 45, 916–925, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Maher et al.(2018)Maher, Matei, Milinski, and
Marotzke</label><mixed-citation>
Maher, N., Matei, D., Milinski, S., and Marotzke, J.: ENSO change in climate
projections: forced response or internal variability?, Geophys. Res. Lett., 45, 11390–11398, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Maher et al.(2019)Maher, Milinski, Suárez-Gutiérrez, Botzet, Dobrynin, Kornblueh, Kröger, Takano, Ghosh, Hedemann, Li, Li, Manzini, Notz, Putrasahan, Boysen, Claussen, Ilyina, Olonscheck, Raddatz, Stevens, and Marotzke</label><mixed-citation>
Maher, N., Milinski, S., Suárez-Gutiérrez, L., Botzet, M., Dobrynin,
M., Kornblueh, L., Kröger, J., Takano, Y., Ghosh, R., Hedemann, C., Li,
C., Li, H., Manzini, E., Notz, D., Putrasahan, D., Boysen, L., Claussen, M.,
Ilyina, T., Olonscheck, D., Raddatz, T., Stevens, B., and Marotzke, J.: The
Max Planck Institute Grand Ensemble: Enabling the Exploration of Climate
System Variability, J. Adv. Model. Earth Syst., 28, 867–920, <a href="https://doi.org/10.1029/2019MS001639" target="_blank">https://doi.org/10.1029/2019MS001639</a>, 2019.
</mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Olonscheck and Notz(2017)</label><mixed-citation>
Olonscheck, D. and Notz, D.: Consistently Estimating Internal Climate
Variability from Climate Model Simulations, J. Climate, 30, 9555–9573, <a href="https://doi.org/10.1175/JCLI-D-16-0428.1" target="_blank">https://doi.org/10.1175/JCLI-D-16-0428.1</a>, 2017.

</mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Pausata et al.(2015)Pausata, Grini, Caballero, Hannachi, and
Seland</label><mixed-citation>
Pausata, F. S. R., Grini, A., Caballero, R., Hannachi, A., and Seland, Ø.:
High-latitude volcanic eruptions in the Norwegian Earth System Model: the
effect of different initial conditions and of the ensemble size, Tellus B, 11, 2050–2069, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Rodgers et al.(2015)Rodgers, Lin, and Frölicher</label><mixed-citation>
Rodgers, K. B., Lin, J., and Frölicher, T. L.: Emergence of multiple ocean ecosystem drivers in a large ensemble suite with an Earth system model, Biogeosciences, 12, 3301–3320, <a href="https://doi.org/10.5194/bg-12-3301-2015" target="_blank">https://doi.org/10.5194/bg-12-3301-2015</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Steinman et al.(2015)Steinman, Frankcombe, Mann, Miller, and
England</label><mixed-citation>
Steinman, B. A., Frankcombe, L. M., Mann, M. E., Miller, S. K., and England,
M. H.: Response to Comment on “Atlantic and Pacific multidecadal oscillations and Northern Hemisphere temperatures”, Science, 350, 1326, <a href="https://doi.org/10.1126/science.aac5208" target="_blank">https://doi.org/10.1126/science.aac5208</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Stevenson et al.(2012)Stevenson, Fox-Kemper, Jochum, Neale, Deser,
and Meehl</label><mixed-citation>
Stevenson, S., Fox-Kemper, B., Jochum, M., Neale, R., Deser, C., and Meehl, G.: Will There Be a Significant Change to El Niño in the Twenty-First
Century?, J. Climate, 25, 2129–2145, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Stolpe et al.(2018)Stolpe, Medhaug, Sedláček, and
Knutti</label><mixed-citation>
Stolpe, M. B., Medhaug, I., Sedláček, J., and Knutti, R.: Multidecadal Variability in Global Surface Temperatures Related to the Atlantic Meridional Overturning Circulation, J. Climate, 31, 2889–2906,
<a href="https://doi.org/10.1175/JCLI-D-17-0444.1" target="_blank">https://doi.org/10.1175/JCLI-D-17-0444.1</a>, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Taylor et al.(2012)Taylor, Stouffer, and Meehl</label><mixed-citation>
Taylor, K. E., Stouffer, R. J., and Meehl, G. A.: An Overview of CMIP5 and the Experiment Design, B. Am. Meteorol. Soc., 93, 485–498, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Thompson et al.(2015)Thompson, Barnes, Deser, Foust, and
Phillips</label><mixed-citation>
Thompson, D. W. J., Barnes, E. A., Deser, C., Foust, W. E., and Phillips, A. S.: Quantifying the Role of Internal Climate Variability in Future Climate Trends, J. Climate, 28, 6443–6456, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>von Känel et al.(2017)von Känel, Frölicher, and
Gruber</label><mixed-citation>
von Känel, L., Frölicher, T. L., and Gruber, N.: Hiatus-like decades
in the absence of equatorial Pacific cooling and accelerated global ocean
heat uptake, Geophys. Res. Lett., 44, 7909–7918, <a href="https://doi.org/10.1002/2017GL073578" target="_blank">https://doi.org/10.1002/2017GL073578</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Wittenberg(2009)</label><mixed-citation>
Wittenberg, A. T.: Are historical records sufficient to constrain ENSO
simulations?, Geophys. Res. Lett., 36, 3–5, 2009.
</mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Zelle et al.(2005)Zelle, Jan van Oldenborgh, Burgers, and
Dijkstra</label><mixed-citation>
Zelle, H., van Oldenborgh, G. J., Burgers, G., and Dijkstra, H.: El Niño
and Greenhouse Warming: Results from Ensemble Simulations with the NCAR CCSM, J. Climate, 18, 4669–4683, <a href="https://doi.org/10.1175/JCLI3574.1" target="_blank">https://doi.org/10.1175/JCLI3574.1</a>, 2005.
</mixed-citation></ref-html>--></article>
