Community-driven fair data management and reproducibility for the whole image-data life cycle
Bibliographic record
Abstract
Poster Nr. 1571 presented at the 2022 ASCB and EMBO Cell Biology Meeting held on December 3-7, 2022 in Washington, DC, USA https://www.ascb.org/cellbio2022/ Abstract Biomedical advances crucially depend on the generation of high-quality Findable, Accessible, Interoperable, and Reproducible (FAIR; 10.1038/sdata.2016.18) datasets. This, in turn, requires the seamless integration of community-specified image documentation practices within the Research Data Management (RDM) processing pipelines required to ensure the execution, tracking, and documentation of the entire life cycle of data from sample preparation to publication (i.e., data provenance). This is important for microscopy, where data interpretation is crucially dependent on easy-to-use RDM software enabling the capture and reporting of knowledge that is collectively termed Image Metadata, and that consists of three key aspects: i) biological context (i.e., organism, growth conditions, sample-type); ii) image acquisition (i.e., microscope hardware/settings/quality-control); and iii) image processing (i.e., software, analysis steps). To illustrate these points, this presentation will first introduce recently published 4DN-BINA-OME-QUAREP community-driven Image Metadata specifications developed in the context of international bioimaging initiatives (10.1038/s41592-021-01327-9) and how they can be applied to typical light microscopy experiments. This will be followed by a deep dive into the importance of incorporating robust microscopy quality assessment and reporting procedures in the life cycle of light-microscopy data to ensure rigor, reproducibility, and reusability. The discussion will identify key stages in the pathway that includes image data acquisition, management, analysis, and dissemination and provide OMERO-based concrete and practical examples of how open-source tools and protocols developed by an international consortium of community initiatives led by QUality Assessment and REProducibility in Light Microscopy (QUAREP-LiMi), are being utilized in close collaboration with Canada BioImaging, at McGill University and UMass Medical School to capture and report the necessary quality-control metrics and metadata to support the reproducibility and reusability of image-based datasets. Finally, the presentation will also introduce the Micro-Meta App and MethodsJ2 software tools that allow researchers to collect detailed microscope hardware and acquisition settings metadata and generates draft methods text for publication.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.124 | 0.190 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.006 | 0.005 |
| Science and technology studies | 0.004 | 0.004 |
| Scholarly communication | 0.016 | 0.014 |
| Open science | 0.007 | 0.025 |
| Research integrity | 0.005 | 0.007 |
| Insufficient payload (model declined to judge) | 0.016 | 0.024 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".