The Validity/Reliability of Occupational Performance Measurement: A Synthesis of Research Using the Validity Generalization Method
Bibliographic record
Abstract
In this study, we used the Validity Generalization method (Hunter & Schmidt, 2004; Schmidt & Hunter, 1977) to synthesize findings from 18 studies investigating the validity/reliability of a variety of occupational performance assessments. Our objectives were to determine: 1) an overall estimate of the validity/reliability of occupational performance measurement scores; and 2) generalizability of the validity/reliability from research to clinical settings after correction for some attenuating statistical artifacts. To achieve the above two objectives, we computed weighted mean validity/reliability coefficients of studies validating a variety of occupational performance measurement instruments; determined the attenuation of variability of the validity/reliability coefficients by sampling and test criterion measurement reliability errors; and calculated the proportion of the variance of the validity/reliability coefficients accounted for by the attenuating factors. Our sample was comprised of 18 studies in which test-retest, inter-rater, alternate measure, and predictive reliability estimates of occupational performance scores were investigated. The instruments generating scores in the studies included: the Canadian Occupational Performance Measure (COPM); Occupational Performance History Interview (OPHI-II); Assessment of Motor and Process Skills (AMPS); Role Checklist (RC); Assessment of Living Skills and Resources (ALSAR); Australian Therapy Outcome Measures (AusTOMs); Functional Independence Measure (FIM); and Hessel Analogical Reasoning Test (HART) among others. Occupational Performance assessment scores based on self-report were found to have a higher corrected weighted mean validity/reliability coefficient than is typical for instruments in social science research. This was particularly significant in the context of the emphasis on client-centeredness in the current occupational therapy paradigm which encourages collaboration with clients in the assessment and intervention process. When observed variance was corrected for attenuation by sampling and test criterion reliability errors, less than 75% of the variance recommended by Hunter and Schmidt (2004) remained. Our findings indicated that assessment scores based on self-report instruments may be the most reliable/valid. It is not clear whether such validity/reliability can be assumed in clinical conditions that differ from research circumstances. Further meta-analysis is indicated to determine more conclusively such generalizability of validity/reliability of occupational performance assessments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.396 | 0.553 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.006 | 0.006 |
| Bibliometrics | 0.033 | 0.020 |
| Science and technology studies | 0.003 | 0.010 |
| Scholarly communication | 0.010 | 0.011 |
| Open science | 0.003 | 0.009 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".