From Better Monitoring to Better Decisions: Improving Conservation using Community Science
Bibliographic record
Abstract
Understanding and abating threats to biodiversity requires extensive data collection, and yet doing so is logistically challenging and depletes funds that could otherwise be spent on action.Fortunately, we have already accrued vast quantities of data on biodiversity through community science programs, which enlist help from the public for monitoring and research.My research aims to put these data to work, demonstrating how they can help compensate for monitoring biases, fill knowledge gaps, and improve conservation decision making, all while reducing the cost to do so.First, I assess the current use of community science data in peer-reviewed research, examining taxonomic and geographic patterns in the literature.The next project uses data collected through an opportunistic community science dataset to model population trends under similar frameworks to professional monitoring schemes, while accounting for the additional noise and variability present in such datasets.Then, using community science data available across the entire western hemisphere, I examine how species respond to anthropogenic pressures differently over time and space.Next, I investigate how the use of community science data can help redistribute conservation resources from monitoring to action, with ultimately better outcomes for biodiversity.Finally, I discuss common challenges associated with conducting research using community science data, and explore solutions that are applicable to data collected by amateurs and professionals alike.Although large noisy datasets present many analytical challenges, my research shows that using them can directly improve our capacity to make informed decisions and ultimately lead to more effective and efficient conservation.Conservation science is a crisis discipline, and we must bring all possible tools to bear before species are lost for good.The last 5 years have been some of the best of my life.I came somewhat reluctantly to grad school because it seemed like the only option to move forward in my career at the time.What I discovered here was a real passion for research that will shape the rest of my career, and I am so grateful to the many people who helped make that happen.I want to first thank my family for all their support throughout my education.I know that this journey would have been much, much harder without it, and I am grateful every day that I was afforded such opportunities.I have been very privileged to have so many great mentors guiding me over the years.I am extremely grateful to my committee, who have been there from the start and provided valuable feedback and guidance on this thesis.Scott Wilson, thank you for your exceptional mentorship over the years, both with regards to my research and my career.Adam Smith,
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.076 | 0.246 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.014 | 0.016 |
| Science and technology studies | 0.005 | 0.010 |
| Scholarly communication | 0.019 | 0.040 |
| Open science | 0.004 | 0.015 |
| Research integrity | 0.004 | 0.008 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".