The effect of automated landmark identification on morphometric analyses
Bibliographic record
Abstract
Morphometric analysis of anatomical landmarks allows researchers to identify specific morphological differences between natural populations or experimental groups, but manually identifying landmarks is time-consuming. We compare manually and automatically generated adult mouse skull landmarks and subsequent morphometric analyses to elucidate how switching from manual to automated landmarking will impact morphometric analysis results for large mouse (Mus musculus) samples (n = 1205) that represent a wide range of 'normal' phenotypic variation (62 genotypes). Other studies have suggested that the use of automated landmarking methods is feasible, but this study is the first to compare the utility of current automated approaches to manual landmarking for a large dataset that allows the quantification of intra- and inter-strain variation. With this unique sample, we investigated how switching to a non-linear image registration-based automated landmarking method impacts estimated differences in genotype mean shape and shape variance-covariance structure. In addition, we tested whether an initial registration of specimen images to genotype-specific averages improves automatic landmark identification accuracy. Our results indicated that automated landmark placement was significantly different than manual landmark placement but that estimated skull shape covariation was correlated across methods. The addition of a preliminary genotype-specific registration step as part of a two-level procedure did not substantially improve on the accuracy of one-level automatic landmark placement. The landmarks with the lowest automatic landmark accuracy are found in locations with poor image registration alignment. The most serious outliers within morphometric analysis of automated landmarks displayed instances of stochastic image registration error that are likely representative of errors common when applying image registration methods to micro-computed tomography datasets that were initially collected with manual landmarking in mind. Additional efforts during specimen preparation and image acquisition can help reduce the number of registration errors and improve registration results. A reduction in skull shape variance estimates were noted for automated landmarking methods compared with manual landmarking. This partially reflects an underestimation of more extreme genotype shapes and loss of biological signal, but largely represents the fact that automated methods do not suffer from intra-observer landmarking error. For appropriate samples and research questions, our image registration-based automated landmarking method can eliminate the time required for manual landmarking and have a similar power to identify shape differences between inbred mouse genotypes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.034 | 0.133 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".