EEM-PARAFAC-SOM for assessing variation in the quality of dissolved organic matter: simultaneous detection of differences by source and season
Bibliographic record
Abstract
Environmental context Dissolved organic matter (DOM) is a highly diverse mixture of interacting compounds, which plays a key role in environmental processes in aquatic systems. The quality and functionality of DOM are measured using fluorescence spectroscopy, but established data analysis assumes linear behaviour, limiting the effectiveness of characterisation. We apply self-organising maps to fluorescence composition to improve the assessment of DOM quality and behaviour by visualising the interdependent nature of its components. Abstract Self-organising maps (SOMs) were used to sort the excitation–emission matrices (EEMs) of dissolved organic matter (DOM) based on their multivariate ‘fluorescence composition’ (i.e. each parallel factor analysis (PARAFAC) component loading, viz. ‘Fmax’ value was expressed as a proportion of all Fmax values in each EEM). This sorting provided a simultaneous organisation of DOM according to differences in quality along a 125-km stretch of a large boreal river, corresponding with both source and season. The information provided by the SOM-based spatial organisation of samples was also used to assess the likelihood of PARAFAC model overfitting. Changes in fluorescence composition caused by changing salinity were also assessed for multiple sources. Seasonal and source-based differences were readily apparent for the main stem of the river and tributaries, and source-based differences were apparent in both fresh and saline groundwaters. Proportions of humic-like components were positively correlated with the amounts of bog, fen and swamp in tributary watersheds. Proportions of six PARAFAC components were negatively correlated with the proportions of all wetland types, and positively correlated with the proportions of open water and other land cover. Ancient saline groundwaters contained >50 % protein-like DOM. There was no change in DOM quality from upstream to downstream in August or October. Increasing salinity was associated with additional protein-like fluorescence in all sources, but source-based differences were also apparent. The application of SOM to fluorescence composition is highly recommended for assessing and visualising transformations and differences in DOM quality, and relating them to associated properties.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".