MétaCan
Menu
Back to cohort
Record W4399366685 · doi:10.2166/wst.2024.182

A validation workflow for treatment wetland performance data

2024· article· en· W4399366685 on OpenAlexfundno aff
Sophie Guillaume-Ruty, Josep Pueyo‐Ros, Joaquím Comas, Nicolas Forquet

Bibliographic record

VenueWater Science & Technology · 2024
Typearticle
Languageen
FieldEnvironmental Science
TopicConstructed Wetlands for Wastewater Treatment
Canadian institutionsnot available
FundersHorizon 2020 Framework ProgrammeInstitut National de Recherche pour l'Agriculture, l'Alimentation et l'EnvironnementAgence Nationale de la RechercheGeneralitat de CatalunyaEuropean CommissionIndian National Science AcademyCanadian Institute for Advanced Research
KeywordsWorkflowScope (computer science)AdaptabilityComputer scienceSet (abstract data type)Data miningKey (lock)Data scienceQuality (philosophy)Process managementDatabaseEngineering

Abstract

fetched live from OpenAlex

ABSTRACT Treatment wetlands (TWs) effectively remove target pollutants and enhance urban water circularity and resilience. They constitute a prominent solution for urban wastewater treatment, thanks to their adaptability across various types of wastewater, scales and climatic conditions. However, the disparity in TW designs and the focus on a restricted set of variables applicable to research studies impede any comprehensive evaluation and comparison of TW performance. Our study introduces a methodology for data validation, in concurrently establishing a workflow specific to TW. This approach is aimed at defining the scope and relationships within the data, implementing checks and concatenating them into a quality flag, as an initial step towards building reliable statistical models. We underscore the importance of both mobilising comprehensive knowledge and identifying customary, yet implicit, choices intertwined in data processing. As for the application workflow, we collected and analysed data sourced from peer-reviewed papers on horizontal and vertical flow TW. Deficiencies were noted in key data elements like dimensions, concentrations and operational conditions. For the data analysis, relationships are highlighted between variables introduced for modelling purposes. These methodologies and workflows assess the quality of the data, in paving the way towards more dependable statistical models for TW design and implementation.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.098
metaresearch head score (Gemma)0.207
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.098
Threshold uncertainty score0.516

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0980.207
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0020.003
Bibliometrics0.0140.008
Science and technology studies0.0020.002
Scholarly communication0.0080.005
Open science0.0030.005
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0070.005

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.019
GPT teacher head0.252
Teacher spread0.233 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueWater Science & TechnologySame topicConstructed Wetlands for Wastewater TreatmentFrench-language works237,207