Comparison of discriminant function and classification tree analyses for age classification of marmots
Bibliographic record
Abstract
We evaluated the predictive power of two classification techniques, one parametric - discriminant function analysis (DFA) and the other non?parametric - classification and regression tree analysis (CART), in order to provide a non?subjective quantitative method of determining age class in Vancouver Island marmots (Marmota vancouverensis) and hoary marmots (Marmota caligata). For both techniques we used morphological measurements of known?age male and female marmots from two independent population studies to build and test predictive models of age class. Both techniques had high predictive power (69-86%) for both sexes and both species. Overall, the two methods performed identically with 81% correct classification. DFA was marginally better at discriminating among older more challenging age classes compared to CART. However, in our test samples, cases with missing values in any of the discriminant variables were deleted and hence unclassified by DFA, whereas CART used values from closely correlated variables to substitute for the missing values. Therefore, overall, CART performed better (CART 81% vs DFA 76%) because of its ability to classify incomplete cases. Correct classification rates were approximately 10% higher for hoary marmots than for Vancouver Island marmots, a result that could be attributed to different sets of morphological measurements. Zygomatic arch breadth measured in hoary marmots was the most important predictor of age class in both sexes using both classification techniques. We recommend that CART analysis be performed on data?sets with incomplete records and used as a variable screening tool prior to DFA on more complete data?sets.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".