Eddy Covariance Flux Data: Sitting on a Golden Egg
Bibliographic record
Abstract
Data from thousands of past and present eddy covariance flux stations are available across the globe, while multiple hundreds actively operating as individual process-level studies, small flux networks dedicated to specific research goals, and larger national and continental networks with broad ecological and environmental foci. Many flux stations have weather and soil data to help clean, analyze and interpret the fluxes but most do not have optical proximal sensors, do not allow straightforward coupling with remote sensing (drone, aircraft, satellite, etc.) data, and cannot easily be used for validation of remotely sensed products, ecosystem modeling, or upscaling from field to regional levels. The flux source areas themselves (e.g., flux footprints) are typically not defined in the flux datasets, and the time stamps of the fluxes come in a large number of outdated non-trackable formats. Finally, the past ways of the flux data quality control, analysis and interpretation require a participation of micrometeorological expert (or an entire network) with their own custom codes or exceptional skills in using existing software such as MatLab or VB Tools in Excel. These are the key issues effectively preventing a larger environmental research community and remote sensing community from fully utilizing eddy covariance flux data. In 2016-2020, a set of new tools to collect, process, analyze, time- and space- allocate and share time-synchronized flux data from multiple flux stations were developed and deployed globally. These new tools can be effective in solving most or all of the key issues listed above. The fully automated FluxSuite system combines hardware, software and web services, and does not require an expert to run it. It can be incorporated into a new flux station or added to a present station, using a weatherized remotely-accessible microcomputer, SmartFlux3 which utilizes EddyPro software to calculate fully-processed fluxes in near-real-time, alongside biomet data and flux footprints. All data are merged into a single quality-controlled file timed using PTP time protocol. Remote sensing researchers and modelers without actual physical stations can form “virtual networks” of actual stations by collaborating with tower PIs from different physical networks and flux databases. The very latest development in this overall approach is the flux data analysis software, Tovi, designed to seamlessly ingest the data from the flux stations and to allow a non-micrometeorologist to quality control, analyze and interpret the flux data. It allows rapid execution of the QC/QA and data analysis steps which have been time-consuming and complicated in the past, and other data analysis steps virtually not doable in the past, all using interactive and intuitive GUI, including advanced footprint calculations and flux apportioning necessary for remote sensing community; NEE flux partitioning; automated generation specific lists of references for each workflow; etc. This presentation will show how combinations of these new tools are used by major networks, and describe how this approach can be utilized for matching remote sensing and tower data for ground truthing, improve scientific interactions, and promote a better utilization of the eddy covariance flux data by a wider environmental research community.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.018 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.007 | 0.010 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.092 | 0.095 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".