Predicting Sex from Hand Dimensions using Statistical Models: A Cross-Sectional Study of Medical Students
Bibliographic record
Abstract
Determining sex from dismembered body parts is crucial for forensic and medico-legal investigations. Hand anthropometry offers a practical, non-invasive approach, particularly when utilising statistical models. This study aims to assess sexual dimorphism in hand dimensions and evaluate the effectiveness of statistical methods in determining sex from various hand measurements. Data for this study were obtained from a cross-sectional survey conducted from July to December 2021 involving a sample of 150 undergraduate medical students (78 males, 72 females) aged 18 to 24 years at a private medical college in India. Medical students were selected as a relatively homogeneous population to minimize confounding factors such as occupational variation, lifestyle, and health status, thereby increasing the internal validity of the findings. Measurements of hand length, breadth, and palm length for both the left and right hands were taken. Logistic regression was employed to develop models for sex classification. By using various combinations of explanatory variables, three logistic regression models were fitted to predict sex. Among these models, the one with the best fit was selected as the final model for sex prediction. It was observed that the mean scores of male and female respondents differ significantly. All hand dimensions were significantly larger in males (p < 0.001). According to the best-fitting model, the Right Hand Index (RHI) along with height was identified as the most significant predictors of sex. The best-fitted logistic model achieved 90% accuracy with an AUC of 0.93. Hand dimensions, particularly the Right Hand Indices (RHI), are effective predictors of sex and logistic regression provides a reliable method for forensic identification when complete body parts are unavailable, and this model can be utilised for other types of forensic predictions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.014 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".