A clustering approach to determine biophysical provinces and physical drivers of productivity dynamics in a complex coastal sea
Bibliographic record
Abstract
Abstract. The balance between ocean mixing and stratification influences primary productivity through light limitation and nutrient supply in the euphotic ocean. Here, we apply a hierarchical clustering algorithm (Ward's method) to four factors relating to stratification (wind energy, freshwater index, water-column-averaged vertical eddy diffusivity, and halocline depth), as well as to depth-integrated phytoplankton biomass, extracted from a biophysical ocean model of the Salish Sea. Running the clustering algorithm on 4 years of model output, we identify distinct regions of the model domain that exhibit contrasting wind and freshwater input dynamics, as well as regions of varying water-column-averaged vertical eddy diffusivity and halocline depth regimes. The spatial regionalizations in physical variables are similar in all 4 analyzed years. We also find distinct interannually consistent biological zones. In the northern Strait of Georgia and Juan de Fuca Strait, a deeper winter halocline and episodic summer mixing coincide with higher summer diatom abundance, while in the Fraser River stratified central Strait of Georgia, shallower haloclines and stronger summer stratification coincide with summer flagellate abundance. Cluster-based model results and evaluation suggest that the Juan de Fuca Strait supports more biomass than previously thought. Our approach elucidates probable physical mechanisms controlling phytoplankton abundance and composition. It also demonstrates a simple, powerful technique for finding structure in large datasets and determining boundaries of biophysical provinces.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".