MétaCan
Menu
Back to cohort
Record W4322209094 · doi:10.5194/egusphere-egu23-15087

Applications of an advanced clustering tool for EU AQ monitoring network data analysis

2023· preprint· en· W4322209094 on OpenAlexaff
Joana Soares, Christoffer Stoll, Islen Vallejo, Colin Lee, Paul A. Makar, L. Tarrasón

Bibliographic record

Venuenot available
Typepreprint
Languageen
FieldEnvironmental Science
TopicAir Quality Monitoring and Forecasting
Canadian institutionsEnvironment and Climate Change Canada
Fundersnot available
KeywordsCluster analysisData miningAir quality indexContext (archaeology)Representativeness heuristicHierarchical clusteringComputer scienceSimilarity (geometry)Air pollutionGeographyMachine learningStatisticsArtificial intelligenceMathematicsMeteorology

Abstract

fetched live from OpenAlex

Air quality monitoring networks provide invaluable data for studying human health, environmental impacts, and the effects of policy changes. In a European legislative context, the data collected constitutes the basis for reporting air quality status and exceedances under the Ambient Air Quality Directives (AAQD) following specific requirements. Consequently, the network's representativity and ability to accurately assess the air pollution situation in European countries become a key issue. The combined use of models and measurements is currently understood as the most robust way to map the status of air pollution in an area, allowing it to quantify both the spatial and temporal distribution of pollution. This spatial-temporal information can be used to evaluate the representativeness of the monitoring network and support air quality monitoring design using hierarchical clustering techniques.The hierarchical clustering methodology applied in this context can be used as a screening tool to analyse the level of similarity or dissimilarity of the air concentration data (time-series) within a monitoring network. Hierarchical clustering assumes that the data contains a level of (dis)similarity and groups the station records based on the characteristics of the actual data. The advantage of this type of clustering is that it does not require an a priori assumption about how many clusters there might be, but it can become computationally expensive as the number of time-series increases in size. Three dissimilarity metrics are used to establish the level of similarity (or dissimilarity) of the different air quality measurements across the monitoring network: (1) 1-R, where R is the Pearson linear correlation coefficient, (2) the Euclidean distance (EuD), and (3) multiplication of metric (1) and (2). The metric based on correlation assesses dissimilarities associated with the changes in the temporal variations in concentration. The metric based on the EuD assesses dissimilarities based on the magnitude of the concentration over the period analysed. The multiplication of these two metrics (1-R) x EuD assesses time variation and pollution levels correlations, and it has been demonstrated to be the most useful metric for monitoring network optimization.This study presents the MoNET webtool developed based on the hierarchical clustering methodology. This webtool aims to provide an easy solution for member states to quality control the data reported as a tier-2 level check and evaluate the representativeness of the air quality network reporting under the AAQD. Some examples from the ongoing evaluation of the monitoring site classification carried out as a joint exercise under the Forum for Air Quality Modeling (FAIRMODE) and the National Air Quality Reference Laboratories Network (AQUILA) are available to show the usability of the tool. MoNet should be able to identify outliers, i.e., issues with the data or data series with very specific temporal-magnitude profiles, and to distinguish, e.g., pollution regimes within a country and if it resembles the air quality zones required by the AAQD and set by the member states; stations monitoring high-emitting sources; background regimes vs. a local source driving pollution regime in cities.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.005
metaresearch head score (Gemma)0.019
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Software · Consensus signal: none
Teacher disagreement score0.010
Threshold uncertainty score0.026

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0050.019
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0080.008
Science and technology studies0.0010.000
Scholarly communication0.0020.001
Open science0.0010.002
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0040.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.103
GPT teacher head0.363
Teacher spread0.259 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreSoftware

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes1
Has abstractyes

Explore more

Same topicAir Quality Monitoring and ForecastingFrench-language works237,207