Advancing archaeal research through FAIR resource and data sharing, and inclusive community building
Bibliographic record
Abstract
Over the last two decades archaeal research has expanded into a wide-ranging research field, driven by a fairly small research community. Archaea are now recognized as important players in the One-Health approach and expertise on the biology of archaea has become crucial in the study of a broad range of topics and environments, including the host-associated microbiomes, major nutrient cycles, greenhouse gas metabolism, the cell biology and origin of eukaryotes, adaptation of life to extremes, as well as various biotechnological applications. Here, we summarize existing resources and ongoing efforts in the engaged broader archaeal scientific community to accelerate research and resource sharing guided by FAIR (findable, accessible, interoperable, reusable) data-sharing principles. We highlight ongoing community efforts that: (i) aim to share protocols and best practices for working with archaea (e.g. ARCHAEA.bio), (ii) combine large ‘omics datasets for the dissemination of unified, system-wide results (e.g. Archaeal Proteome Project, KBase) and (iii) provide opportunities for scientists to present their work in a supportive environment and to forge connections and collaborations (e.g. Archaea Power Hour). Together, these resources and projects promise to spur and cross-fertilize research, making archaeal research more accessible to a broader and more diverse audience. Archaeal research and its growing importance have benefited from a community that is engaged in various collaborative efforts, which are highlighted here with examples for the sharing of resources and data, as well as the building of supportive environments to advance the community.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.121 | 0.173 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.010 | 0.009 |
| Science and technology studies | 0.009 | 0.013 |
| Scholarly communication | 0.022 | 0.036 |
| Open science | 0.009 | 0.083 |
| Research integrity | 0.005 | 0.006 |
| Insufficient payload (model declined to judge) | 0.009 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".