Training General Movements Classifiers with Global Labels Offers Insights on Sub-movement Quality
Bibliographic record
Abstract
Abnormal or absent General Movements (GMs) during the fidgety period (9-20 weeks post-term) are strong early indicators of neurological disorders, including cerebral palsy (CP).The General Movements Assessment (GMA) is the clinical gold standard for evaluating GMs, but its reliance on expert assessment limits accessibility and scalability.Machine Learning (ML)-based models offer a promising alternative in automating movement classification; however, existing approaches to training these systems either require extensive manual annotation of short segments of infant movements (or snippets) or classify entire videos without capturing movement-level details.This study addresses these limitations by demonstrating that a ML classifier trained with video-level (per infant) labels can accurately classify whole videos of infant movements and provide useful information about movement snippets.We trained and evaluated several models, including SVM, LSTM, 1D-CNN, and Vision Transformer (ViT), using time-series representations of infant movements.The best-performing model, a 1D-CNN, achieved 100% accuracy in video-level classification and 87.5% accuracy in snippet-level (i.e., movement-level) classification of previously unseen data, using 2D coordinates of 24 body landmarks and 12 joint angle features.Additionally, we examined whether the feature space of videos labelled as normal and abnormal shows overlap, using Independent Component Analysis (ICA) and cosine-similarity between 1D-CNN abstractions.Our findings align with clinical observations, indicating that short movement segments from infants labelled as abnormal share characteristics with those from infants with normal GMs, impacting classification performance.Overall, this work provides insights useful for working towards fully automated GMA analysis capable of providing both movement-and video-level assessment, which will enhance early prediction of neurodevelopmental abnormalities with improved scalability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".