Comment on “Evaluation of surface albedo and snow cover in AR4 coupled models” by A. Roesch
Bibliographic record
Abstract
[1] We believe that one of the main conclusions reached by Roesch [2006, paragraph 1] (hereinafter referred to as R2006) from his evaluation of the performance of the Intergovernmental Panel on Climate Change (IPCC) Fourth Assessment Review (AR4) climate model snow cover simulations, i.e., that "most IPCC AR4 climate models predict excessive snow mass in spring and suffer from a delayed spring snow melt …", is incorrect for North America and probably for Eurasia as well. This erroneous conclusion is a result of deficiencies in the U.S. Air Force (USAF) global snow depth climatology of Foster and Davy [1988] (hereinafter referred to as USAF) which were previously documented by Brown et al. [2003] (hereinafter referred to as B2003) and is exacerbated by the application of an inappropriate method for estimating snow-water equivalent (SWE) from snow depth. [2] R2006 (paragraph 13) stated that the USAF data set was "generally considered to be one of the most reliable and accurate snow depth climatologies available." This statement is supported neither by the conclusions of B2003 nor by the USAF data set documentation. The USAF data set is a hand-drawn analysis based on manual interpretation of a wide range of mainly hard copy sources (e.g., atlases and climate summaries) covering various time periods. Foster and Davy [1988] clearly indicate in the documentation that they had varying levels of confidence in their data sources with low confidence in most of the high-latitude data. In mountainous areas, Foster and Davy [1988] used mean monthly precipitation totals from World Meteorological Organization climatic atlases (1970 and 1981) and assumed a mean snow density of 300 kg m−3 to arrive at a snow depth. The manual analysis was digitized on a 47 km Northern Hemisphere polar stereographic (PS) grid, and it is not clear where R2006 obtained the 1° × 1° equal-area grid version of the data. In the comparisons included in this comment the USAF data were used in their original gridded format. [3] B2003 carried out a detailed daily snow depth analysis on a 1/3° latitude-longitude grid over North America (NA) for the second phase of the Atmospheric Model Intercomparison Project (AMIP-2) period (1979–1996) that incorporated ∼8000 snow depth observations per day from U.S. cooperative stations and Canadian climate stations. The snow depth analysis is based on the scheme developed by Brasnett [1999] and employed operationally at the Canadian Meteorological Centre. The first-guess field used a simple snow accumulation, aging, and melt model driven by 6-hourly values of air temperature and precipitation from the European Centre for Medium-Range Weather Forecasting Reanalysis, December 1978 to February 1994 (ERA-15) with extensions from the Tropical Ocean–Global Atmosphere operational data archive. The gridded snow depth and estimated SWE results were found to agree well with available independent in situ and satellite data over midlatitudinal regions of the continent, and the snow depth climatology exhibited several improvements over USAF, namely, a more realistic representation of snow depths over the western cordillera and improved snow cover extent in the fall and spring periods when compared to the NOAA weekly snow cover product [Robinson et al., 1993]. The problems with the USAF product during the spring melt period are highlighted in Figure 1, which demonstrates that it ablates the snowpack too rapidly in the spring, especially over the western cordillera, resulting in significant underestimations of SWE as well as snow-covered area. For the comparisons included in this paper the B2003 data set was reinterpolated to the same 47 km PS grid as the USAF data set. [4] The B2003 data set was used to evaluate snow cover simulations of the atmospheric general circulation models participating in the AMIP-2 project by Frei et al. [2005] (several of these models were included in the R2006 evaluation), and the conclusion, contrary to R2006, was that most of the models simulated the seasonal timing and relative spatial patterns of continental-scale SWE fairly well, with a tendency to overestimate the rate of ablation during spring (compare Figure 1 of R2006 to Figure 1a of Frei et al. [2005]). This tendency can also be observed in AMIP-2–simulated snow cover extent for NA [Frei et al., 2003, Figure 3a]. It will now be shown that the different conclusions are attributable to (1) underestimated snow depths over NA during the spring period in the USAF data set and (2) the SWE-density relationship used by R2006 which systematically underestimates density in the spring period. [5] Figure 2 compares monthly mean snow depths from the USAF and B2003 climatologies. The monthly values represent spatially averaged means over the entire NA landmass north of 20°N. The two data sets are in reasonable agreement from November to March, but the USAF displays large underestimates of snow depth in April and May related primarily to overly rapid ablation of snow cover over the western cordillera (see Figure 1). [7] The excessive spring ablation in the USAF data set over NA and the unrealistic density expression combine to produce a serious underestimate of the April-May SWE (Figure 4) on the order of 30 mm which represents about 50% of the estimated April SWE and 75% of the estimated May SWE. The B2003 SWE for NA falls in the middle of the model results shown by R2006 (his Figure 1c), which confirms the findings of Frei et al. [2005] that the current generation of climate models are able to provide realistic simulations of the continental-scale SWE over NA. [8] There is no detailed snow depth or SWE reanalysis available for evaluating the R2006 SWE results for Eurasia. However, it is possible to obtain some insight into systematic biases in the USAF data set by comparing snow cover extent (SCE) with the SCE climatology from the NOAA weekly data set. SCE was estimated from snow depth in the USAF data set using an empirically derived expression from B2003 (their equation (11)) relating in situ snow depth to the snow cover extent seen by the NOAA snow cover product. The results (Figure 5) show that USAF systematically underestimates SCE by about 10–15% during the period of November–March and by 30–50% in April and May. This underestimation in SCE combined with the R2006 underestimation of spring snow cover density helps explain why the USAF SWE climatology for Eurasia is lower than nearly all of the AR4 climate models. Computing SWE from the USAF data set with a fixed density of 300 kg m−3 (consistent with the density used by Foster and Davy [1988] to estimate depths in mountainous regions with the highest SWE values) and adjusting for the SCE bias observed in the NOAA comparison yields average SWE values for Eurasia that are about 15 mm higher than R2006 over the December–April period, which places the USAF results in the middle of the AR4 models results for Eurasia (R2006, Figure 1b). [9] Jim Foster (NASA/GSFC) is acknowledged for providing the USAF snow depth climatology in its original gridded format. David Robinson (Rutgers University) is acknowledged for providing monthly values of continental snow cover extent from the quality-controlled version of the NOAA weekly data set maintained at Rutgers University.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.065 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.005 | 0.005 |
| Scholarly communication | 0.004 | 0.007 |
| Open science | 0.006 | 0.003 |
| Research integrity | 0.039 | 0.041 |
| Insufficient payload (model declined to judge) | 0.008 | 0.012 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".