Principle 5: Sub-select data

Important Considerations
  • For some regions, there may be published analyses or distributions of models that can be leveraged for sampling (e.g. distribution plots characterizing models from hot to cold or wet to dry are often included within credibility assessments like Krantz et al. 2021).
  • Sub-selection always involves tradeoffs, so it is important to do so carefully and transparently. Even with the most careful sampling procedure, there may be a chance of missing plausible worst-case scenarios.
  • When sub-selecting, it is essential to avoid unintentionally selecting a biased sample (e.g. unknowingly choosing the hottest three models from a larger set).
  • Using the entire sample of all credible models will provide the most comprehensive picture of future projections. But if sub-selecting is necessary, most scientific recommendations suggest that at a minimum the sample should contain 5-10 models, although there may be differences based on specific variables, size of change signal, etc. (see Pierce et al. 2009, McSweeney et al. 2014, Mendlik and Gobiet 2015, and CA DWR 2015).
  • Sub-selection strategies may differ for applications that require extreme event analysis and internal variability, versus those that do not require analyzing extreme event statistics or internal variability.
    • Applications involving the evaluation of statistics of extreme events or internal variability may require large ensembles to ensure that there are sufficient instances of rare events to adequately examine their properties and frequencies. Therefore these applications may rely on a suite of models with multiple ensemble members.
    • For applications that do not require these analyses, it may not be essential to distinguish among models and ensembles, and each projection can be considered a realization of a plausible outcome.
  • For some specific applications like stress tests, it may not be necessary to examine the entire range of plausible outcomes, and users might specifically choose the models or simulations that fall in the tails of the distribution.