Should We Stop Using Lexical Diversity Measures in Children's Language Sample Analysis?
Bibliographic record
Abstract
PURPOSE: Prior work has identified weaknesses in commonly used indices of lexical diversity in spoken language samples, such as type-token ratio (TTR) due to sample size and elicitation variation, we explored whether TTR and other diversity measures, such as number of different words/100 (NDW), vocabulary diversity (VocD), and the moving average TTR would be more sensitive to child age and clinical status (typically developing [TD] or developmental language disorder [DLD]) if samples were obtained from standardized prompts. METHOD: We utilized archival data from the norming samples of the Test of Narrative Language and the Edmonton Narrative Norms Instrument. We examined lexical diversity and other linguistic properties of the samples, from a total of 1,048 children, ages 4-11 years; 798 of these were considered TD, whereas 250 were categorized as having a language learning disorder. RESULTS: TTR was the least sensitive to child age or diagnostic group, with good potential to misidentify children with DLD as TD and TD children as having DLD. Growth slopes of NDW were shallow and not very sensitive to diagnostic grouping. The strongest performing measure was VocD. Mean length of utterance, TNW, and verbs/utterance did show both good growth trajectories and ability to distinguish between clinical and typical samples. CONCLUSIONS: This study, the largest and best controlled to date, re-affirms that TTR should not be used in clinical decision making with children. A second popular measure, NDW, is not measurably stronger in terms of its psychometric properties. Because the most sensitive measure of lexical diversity, VocD, is unlikely to gain popularity because of reliance on computer-assisted analysis, we suggest alternatives for the appraisal of children's expressive vocabulary skill.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".