Modelling Riverine Dissolved Silica on Different Spatial and Temporal Scales using Statistical and Machine Learning Methods
Bibliographic record
Abstract
Changes in riverine delivery of dissolved silica (DSi) from continents to the ocean have consequences for the global carbon cycle and marine ecosystem health because (i) rivers are the main source of DSi to the oceans, (ii) diatoms, photosynthetic algae, depend on DSi to build their siliceous frustules, (iii) diatoms are an important foundation of many marine food webs, and (iv) diatoms dominate marine primary productivity and carbon export to the deep ocean. However, despite its importance, the controls of river DSi export are not fully constrained and there is no satisfactory model of riverine DSi yield (DSiY). This thesis investigates river DSiY using statistical and machine learning (ML) methods at different spatial and temporal scales. First, relationships between DSiY and environmental variables were analysed in western Canada to explore the controls of river DSiY on a regional scale. Statistical analyses indicated that land cover and climate are the best predictors of river DSiY in western Canada. Second, a ML model was developed to predict global river DSiY using 30 environmental variables including temperature, precipitation, land cover, lithology, and terrain. Of eight ML algorithms tested, one-step boosted random forest was the best method to predict river DSiY and climate and land cover were the most important predictors in the model, indicating that vegetation affects river DSiY on a global scale. Third, the ML model was used to predict changes in global river DSiY for the future (2090-2099, 2190-2199, and 2290- 2299) and mid-Holocene (6000 years ago). Global mean DSiY is projected to decrease in the future, but predictions have high spatial variability. And, global mean mid-Holocene river DSiY was estimated to be 30% lower than present-day, indicating that changes in paleo-river DSiY may be underestimated. This work demonstrates that DSiY is not only affected by climate and land cover, but that changes in climate and land cover have the potential to influence DSiY on different time scales. This thesis contributes the first study to use ML to model river DSiY and uses a more exhaustive set of variables than other studies, particularly a wider range of land cover variables.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".