MétaCan
Menu
Back to cohort
Record W270785902 · doi:10.1016/j.ecss.2015.05.024

Increasing the quality, comparability and accessibility of phytoplankton species composition time-series data

2015· article· en· W270785902 on OpenAlexaff
Adriana Zingone, Paul J. Harrison, Alexandra Kraberg, Sirpa Lehtinen, Abigail McQuatters‐Gollop, Todd O’Brien, Jun Sun, Hans Jakobsen

Bibliographic record

VenueEstuarine Coastal and Shelf Science · 2015
Typearticle
Languageen
FieldEarth and Planetary Sciences
TopicMarine Biology and Ecology Research
Canadian institutionsUniversity of British Columbia
FundersNational Natural Science Foundation of ChinaNatural Environment Research CouncilSight Research UK
KeywordsPhytoplanktonComparabilityMetadataSampling (signal processing)Computer scienceData qualityEnvironmental scienceRange (aeronautics)HarmonizationData collectionEcologyStatisticsBiologyMathematicsWorld Wide WebEngineeringTelecommunications

Abstract

fetched live from OpenAlex

Phytoplankton diversity and its variation over an extended time scale can provide answers to a wide range of questions relevant to societal needs. These include human health, the safe and sustained use of marine resources and the ecological status of the marine environment, including long-term changes under the impact of multiple stressors. The analysis of phytoplankton data collected at the same place over time, as well as the comparison among different sampling sites, provide key information for assessing environmental change, and evaluating new actions that must be made to reduce human induced pressures on the environment. To achieve these aims, phytoplankton data may be used several decades later by users that have not participated in their production, including automatic data retrieval and analysis. The methods used in phytoplankton species analysis vary widely among research and monitoring groups, while quality control procedures have not been implemented in most cases. Here we highlight some of the main differences in the sampling and analytical procedures applied to phytoplankton analysis and identify critical steps that are required to improve the quality and inter-comparability of data obtained at different sites and/or times. Harmonization of methods may not be a realistic goal, considering the wide range of purposes of phytoplankton time-series data collection. However, we propose that more consistent and detailed metadata and complementary information be recorded and made available along with phytoplankton time-series datasets, including description of the procedures and elements allowing for a quality control of the data. To keep up with the progress in taxonomic research, there is a need for continued training of taxonomists, and for supporting and complementing existing web resources, in order to allow a constant upgrade of knowledge in phytoplankton classification and identification. Efforts towards the improvement of metadata recording, data annotation and quality control procedures will ensure the internal consistency of phytoplankton time series and facilitate their comparability and accessibility, thus strongly increasing the value of the precious information they provide. Ultimately, the sharing of quality controlled data will allow one to recoup the high cost of obtaining the data through the multiple use of the time-series data in various projects over many decades.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.064
metaresearch head score (Gemma)0.190
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Open science
Consensus categoriesnone
DomainCandidate signal: Reproducibility · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.997
Threshold uncertainty score0.341

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0640.190
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0110.013
Science and technology studies0.0010.002
Scholarly communication0.0070.007
Open science0.0030.006
Research integrity0.0020.003
Insufficient payload (model declined to judge)0.0040.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.094
GPT teacher head0.325
Teacher spread0.231 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designTheoretical or conceptual
DomainReproducibility
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations48
Published2015
Admission routes1
Has abstractyes

Explore more

Same venueEstuarine Coastal and Shelf ScienceSame topicMarine Biology and Ecology ResearchFrench-language works237,207