FRESHWATER DIATOMS FROM THE CANADIAN ARCTIC TREELINE AND DEVELOPMENT OF PALEOLIMNOLOGICAL INFERENCE MODELS<sup>1</sup>
Bibliographic record
Abstract
Relationships between surface sediment diatom assemblages and measured environmental variables from 77 lakes in the central Canadian arctic treeline region were examined using multivariate statistical methods. Lakes were distributed across the arctic treeline from boreal forest to arctic tundra ecozones, along steep climatic and environmental gradients. Forward selection in canonical correspondence analysis determined that dissolved inorganic carbon (DIC), dissolved organic carbon (DOC), total nitrogen (TN), lake surface area, silica, lake‐water depth, and iron explained significant portions of diatom species variation. Weighted‐averaging (WA) regression and calibration techniques were used to develop inference models for DIC, DOC, and TN from the estimated optima of the diatom taxa to these environmental variables. Simple WA models with classical deshrinking produced models with the strongest predictive abilities for all three variables based on the bootstrapped root mean squared errors of prediction (RMSEP). WA partial least squares showed little improvement over the simpler WA models as judged by the jackknifed RMSEP. These models suggest that it is possible to infer trends in DIC, DOC, and TN from fossil diatom assemblages from suitably chosen lakes in the central Canadian arctic treeline region.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".