A Registration and Deep Learning Approach to Automated Landmark Detection for Geometric Morphometrics
Bibliographic record
Abstract
ABSTRACT Geometric morphometrics is the statistical analysis of landmark-based shape variation and its covariation with other variables. Over the past two decades, the gold standard of landmark data acquisition has been manual detection by a single observer. This approach has proven accurate and reliable in small-scale investigations. However, big data initiatives are increasingly common in biology and morphometrics. This requires fast, automated, and standardized data collection. Image registration, or the spatial alignment of images, is a fundamental technique in automatic image analysis that is well-poised for such purposes. Yet, in the few studies that have explored the utility of registration-based landmarks for geometric morphometrics, relatively high or catastrophic labelling errors around anatomical extrema are common. Such errors can result in misleading representations of the mean shape, an underestimation of biological signal, and altered variance-covariance patterns. We combine image registration with a deep and domain-specific neural network to automate and optimize anatomical landmark detection for geometric morphometrics. Using micro-computed tomography images of genetically and morphologically variable mouse skulls, we test our landmarking approach under a variety of registration conditions, including different non-linear deformation frameworks (small vs. large) and atlas strategies (single vs. multi). Compared to landmarks derived from conventional image registration workflows, our optimized landmark data show significant reductions in error at problematic locations (up to 0.63 mm), a 36.4% reduction in average landmark coordinate error, and up to a 45.1% reduction in total landmark distribution error. We achieve significant improvements in estimates of the sample mean shape and variance-covariance structure. For biological imaging datasets and morphometric research questions, our method can eliminate the time and subjectivity of manual landmark detection whilst retaining the biological integrity of these expert annotations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".