Structure and representation of ecological data to support knowledge discovery: A case study with bioacoustic data
Bibliographic record
Abstract
Bird communities have long been surveyed as key indicators of ecosystem health and biodiversity. Adoption of Autonomous Recording Units (ARUs) to perform avian surveys has shifted the burden of species recognition from “birders” in the field, to “listeners” who review the ARU recordings at a later time. The number of recordings ARUs can produce has created a need to process large amounts of data. Although much research is devoted to fully automating the recognition process, expert humans are still required when entire bird communities must be identified. A framework for a Decision Support System (DSS) is presented which would assist listeners by suggesting likely species. A unique feature of the DSS is the consideration of the recording “context” of time, location and habitat as well as the bioacoustic features to match unknown vocalizations with reference species. In this thesis a data warehouse was built for an existing set of bioacoustic research data as a first–step to creating the DSS. The data set was from ARU deployments in the Lower Athabasca Region of Alberta, Canada. The Knowledge Discovery in Databases (KDD) and Dimensional Design Process protocols were used as guides to build a Kimball–style data warehouse. Data housed in the data warehouse included field data, data derived from GIS analysis, fuzzy logic memberships and symbolic representation of bioacoustic recording using the Piecewise Aggregate Approximation and Symbolic Aggregate approXimation (PAA/SAX). Examples of how missing and erroneous data were detected and processed are given. The sources of uncertainty inherent in ecological data are discussed and fuzzy logic is demonstrated as a soft–computing technique to accommodate this data. Data warehouses are commonly used for business applications but are very applicable for ecological data. As most instructions on building data warehouse are for business data, this thesis is offered as an example for ecologists interested in moving their data to a data warehouse. This thesis presents a case–study of how a data warehouse can be constructed for existing ecological data, whether as part of a DSS or a tool for viewing research data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.035 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.008 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.007 | 0.007 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".