The effect of automated landmark identification on morphometric analyses
Bibliographic record
Abstract
Morphometric analysis of anatomical landmarks allows researchers to identify specific morphological differences between natural populations or experimental groups, but manually identifying landmarks is time-consuming. We compare manually and automatically generated adult mouse skull landmarks and subsequent morphometric analyses to elucidate how switching from manual to automated landmarking will impact morphometric analysis results for large mouse (Mus musculus) samples (n = 1205) that represent a wide range of 'normal' phenotypic variation (62 genotypes). Other studies have suggested that the use of automated landmarking methods is feasible, but this study is the first to compare the utility of current automated approaches to manual landmarking for a large dataset that allows the quantification of intra- and inter-strain variation. With this unique sample, we investigated how switching to a non-linear image registration-based automated landmarking method impacts estimated differences in genotype mean shape and shape variance-covariance structure. In addition, we tested whether an initial registration of specimen images to genotype-specific averages improves automatic landmark identification accuracy. Our results indicated that automated landmark placement was significantly different than manual landmark placement but that estimated skull shape covariation was correlated across methods. The addition of a preliminary genotype-specific registration step as part of a two-level procedure did not substantially improve on the accuracy of one-level automatic landmark placement. The landmarks with the lowest automatic landmark accuracy are found in locations with poor image registration alignment. The most serious outliers within morphometric analysis of automated landmarks displayed instances of stochastic image registration error that are likely representative of errors common when applying image registration methods to micro-computed tomography datasets that were initially collected with manual landmarking in mind. Additional efforts during specimen preparation and image acquisition can help reduce the number of registration errors and improve registration results. A reduction in skull shape variance estimates were noted for automated landmarking methods compared with manual landmarking. This partially reflects an underestimation of more extreme genotype shapes and loss of biological signal, but largely represents the fact that automated methods do not suffer from intra-observer landmarking error. For appropriate samples and research questions, our image registration-based automated landmarking method can eliminate the time required for manual landmarking and have a similar power to identify shape differences between inbred mouse genotypes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".