Governance and stewardship for research data and information sharing: Issues and prospective solutions in the transdisciplinary plant phenotyping and imaging research center network
Bibliographic record
Abstract
Societal Impact Statement High‐throughput plant phenotyping is a transdisciplinary field of research that provides a systematic approach to assessing and understanding the life cycle of plants and aims to integrate plant genotype with plant ecophysiology and agronomy. Sharing data and information (D&I) is key to accelerating new scientific breakthroughs and innovations in plant phenotyping. The development of a governance and stewardship framework can overcome institutional, legal, and technical barriers that limit D&I sharing. This framework governs how decisions are made about D&I, how researchers engage with each other to manage D&I, and improves the content, discoverability, accessibility, and usability of D&I from different sources. Summary Despite the widely acknowledged value to be added by sharing research data and information (D&I), significant institutional, legal, and technical barriers remain to be addressed, particularly in managing transdisciplinary research. The Plant Phenotyping and Imaging Research Centre (P2IRC) offers an illuminating case study in how various researchers are exchanging their D&I to accommodate the requirements of transdisciplinary research. Using social network analysis to explore the current state of D&I sharing in P2IRC, we found that D&I sharing was modest, with most exchanges taking place between academic researchers and predominantly within the same discipline. Lacking benefits arising from broader dissemination, scholars are not optimally sharing their D&I. Many researchers identified a range of barriers that limited sharing. These barriers are mainly rooted in the absence of D&I governance and stewardship. To overcome barriers, we offer recommendations and solutions for developing a governance and stewardship framework that is tailored to the P2IRC project, but relevant to the wider research world. This includes how to assist adoption of a governance model, accountability and oversight mechanisms, and an operational framework for effective data collection, analysis, possession, and dissemination. We provide recommendations that enable data stewards to preserve and improve the content, discoverability accessibility, and usability of data and metadata by making data interoperable from semantic, syntactic/technical, and legal perspectives. A governance framework would help open all aspects of research outputs, while achieving the access and benefit‐sharing (ABS) goal and contributing to the preservation and sustainable use of D&I.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.137 | 0.202 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.017 | 0.021 |
| Scholarly communication | 0.031 | 0.029 |
| Open science | 0.006 | 0.027 |
| Research integrity | 0.008 | 0.011 |
| Insufficient payload (model declined to judge) | 0.007 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".