Data download speed test for CMIP6 model output: preliminary results
Bibliographic record
Abstract
The World Climate Research Programme (WCRP) facilitates analysis and prediction of Earth system change for use in a range of practical applications of direct relevance, benefit and value to society. WCRP initialized the Coupled Model Intercomparison Project (CMIP) in 1995. The aim of CMIP is to better understand past, present and future climate changes arising from natural, unforced variability or in response to changes in radiative forcing in a multi-model context. The climate model output data that are being produced during this sixth phase of CMIP (CMIP6) is expected to be 40~60 PB. It is still not very clear whether researchers worldwide may experience a big problem when downloading such a huge volume of data. This work addressed this issue by performing data download speed test for all the CMIP6 data nodes. A Google Chrome-based data download speed test website (http://speedtest.theropod.tk) was implemented. It leverages the Allow CORS: Access-Control-Allow-Origin extension to access to each CMIP6 data node. This test consists of four steps: Installing and enabling Allow CORS extension in Chrome, performing data download speed test for all the CMIP6 data nodes, presenting the test results, and uninstalling the extension. The speed test is performed by downloading a certain chunk of model output data file from the thredds data server of each data node. Researchers from 11 countries have performed this test in 24 cities against all the 26 CMIP6 data nodes. The fastest transfer speed was 124MB/s, and the slowest were 0 MB/s because of connect timeout. Data transfer speed in developed countries (United States, Netherland, Japan, Canada, Great Britain) is significantly faster than that in developing countries (China, India, Russia, Pakistan). In developed countries the data transfer mean speed is roughly 80Mb/s, equal to the median US residential broadband speed provided by cable or fiber(FCC Measuring Fixed Broadband - Eighth Report, but in developing countries the mean transfer speed is usually much slower, roughly 9Mb/s. Data transfer speed was significantly faster when the data nodes and test sites were both at developed countries, for example, downloading data from IPSL, DKRZ or GFDL at Wolvercote, UK. Although further test are definitely needed, this preliminary result clearly show that the actual data download speed varies dramatically in different countries, and for different data node. This suggests that ensuring smooth access to CMIP6 data is still challenging.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.017 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.013 | 0.007 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".