Fundamental research questions in subterranean biology
Bibliographic record
Abstract
Five decades ago, a landmark paper in Science titled The Cave Environment heralded caves as ideal natural experimental laboratories in which to develop and address general questions in geology, ecology, biogeography, and evolutionary biology. Although the 'caves as laboratory' paradigm has since been advocated by subterranean biologists, there are few examples of studies that successfully translated their results into general principles. The contemporary era of big data, modelling tools, and revolutionary advances in genetics and (meta)genomics provides an opportunity to revisit unresolved questions and challenges, as well as examine promising new avenues of research in subterranean biology. Accordingly, we have developed a roadmap to guide future research endeavours in subterranean biology by adapting a well-established methodology of 'horizon scanning' to identify the highest priority research questions across six subject areas. Based on the expert opinion of 30 scientists from around the globe with complementary expertise and of different academic ages, we assembled an initial list of 258 fundamental questions concentrating on macroecology and microbial ecology, adaptation, evolution, and conservation. Subsequently, through online surveys, 130 subterranean biologists with various backgrounds assisted us in reducing our list to 50 top-priority questions. These research questions are broad in scope and ready to be addressed in the next decade. We believe this exercise will stimulate research towards a deeper understanding of subterranean biology and foster hypothesis-driven studies likely to resonate broadly from the traditional boundaries of this field.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.016 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.003 | 0.023 |
| Scholarly communication | 0.007 | 0.020 |
| Open science | 0.002 | 0.005 |
| Research integrity | 0.005 | 0.008 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".