Comparison of ocean-colour algorithms for particulate organic carbon in global ocean
Bibliographic record
Abstract
In the oceanic surface layer, particulate organic carbon (POC) constitutes the biggest pool of particulate material of biological origin, encompassing phytoplankton, zooplankton, bacteria, and organic detritus. POC is of general interest in studies of biologically-mediated fluxes of carbon in the ocean, and over the years, several empirical algorithms have been proposed to retrieve POC concentrations from satellite products. These algorithms can be categorised into those that make use of remote-sensing-reflectance data directly, and those that are dependent on chlorophyll concentration and particle backscattering coefficient derived from reflectance values. In this study, a global database of in situ measurements of POC is assembled, against which these different types of algorithms are tested using daily matchup data extracted from the Ocean Colour Climate Change Initiative (OC-CCI; version 5). Through analyses of residuals, pixel-by-pixel uncertainties, and validation based on optical water types, areas for POC algorithm improvement are identified, particularly in regions underrepresented in the in situ POC data sets, such as coastal and high-latitude waters. We conclude that POC algorithms have reached a state of maturity and further improvements can be sought in blending algorithms for different optical water types when the required in situ data becomes available. The best performing band ratio algorithm was tuned to the OC-CCI version 5 product and used to produce a global time series of POC between 1997–2020 that is freely available.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".