Harnessing Deep Learning to Analyze Cryptic Morphological Variability of <i>Marchantia polymorpha</i>
Bibliographic record
Abstract
Characterizing phenotypes is a fundamental aspect of biological sciences, although it can be challenging due to various factors. For instance, the liverwort Marchantia polymorpha is a model system for plant biology and exhibits morphological variability, making it difficult to identify and quantify distinct phenotypic features using objective measures. To address this issue, we utilized a deep-learning-based image classifier that can handle plant images directly without manual extraction of phenotypic features and analyzed pictures of M. polymorpha. This dioicous plant species exhibits morphological differences between male and female wild accessions at an early stage of gemmaling growth, although it remains elusive whether the differences are attributable to sex chromosomes. To isolate the effects of sex chromosomes from autosomal polymorphisms, we established a male and female set of recombinant inbred lines (RILs) from a set of male and female wild accessions. We then trained deep learning models to classify the sexes of the RILs and the wild accessions. Our results showed that the trained classifiers accurately classified male and female gemmalings of wild accessions in the first week of growth, confirming the intuition of researchers in a reproducible and objective manner. In contrast, the RILs were less distinguishable, indicating that the differences between the parental wild accessions arose from autosomal variations. Furthermore, we validated our trained models by an 'eXplainable AI' technique that highlights image regions relevant to the classification. Our findings demonstrate that the classifier-based approach provides a powerful tool for analyzing plant species that lack standardized phenotyping metrics.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".