Predicting coral mortality in South East Asia using open-source data
Bibliographic record
Abstract
Coral reefs hold some of the most biodiverse and productive environments in the world while providing goods and services across a wide range of sectors. Unfortunately, due to climate change and other anthropogenic factors coral reefs are dying. However, majority of reef modeling literature looks at coral development rather than its mortality. In this study I modeled coral mortality using exclusively open source data and created a suitability score (concern index) for high concern corals that also considers the economic value of reefs. The goal was to help identify which community management units require attention from the marine conservation charity Blue Ventures. An inventory of global resolution datasets including environmental and management practices that may indicate coral stress were added to an empty linear model to predict mortality in South East Asia. The best model in predicting coral mortality utilized artisanal fishing data, the model had an R2 of 0.8673 and an accuracy of 48.51%. This model identified the locations of the highest mortality to be in Timor-Leste at Blue Ventures management unit sites TL1 and TL5, however adding economic value in the concern index then Papua New Guinea site PNG7 is more at risk. I hypothesize that a lack in variability and standardized approach in the coral mortality/bleaching explanatory data decreases the accuracy and real-world applicability of the model. The majority of explanatory variables were obtained from the online database Reef Base, some of the observations showed extreme outliers which were all indicated by citizen scientists. Going forward I would recommend that further education and standardization is included in observations to ensure the open source data has a higher accuracy to make better models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".