Validation of Administrative Osteoarthritis Diagnosis Using a Clinical and Radiological Population-Based Cohort
Bibliographic record
Abstract
Objectives.The validity of administrative osteoarthritis (OA) diagnosis in British Columbia, Canada, was examined against X-rays, magnetic resonance imaging (MRI), self-report, and the American College of Rheumatology criteria.Methods.During 2002–2005, 171 randomly selected subjects with knee pain aged 40–79 years underwent clinical assessment for OA in the knee, hip, and hands. Their administrative health records were linked during 1991–2004, in which OA was defined in two ways: (AOA1) at least one physician’s diagnosis or hospital admission and (AOA2) at least two physician’s diagnoses in two years or one hospital admission. Sensitivity, specificity, and predictive values were compared using four reference standards.Results.The mean age was 59 years and 51% were men. The proportion of OA varied from 56.3 to 89.7% among men and 77.4 to 96.4% among women according to reference standards. Sensitivity and specificity varied from 21 to 57% and 75 to 100%, respectively, and PPVs varied from 82 to 100%. For MRI assessment, the PPV of AOA2 was 100%. Higher sensitivity was observed in AOA1 than AOA2 and the reverse was true for specificity and PPV.Conclusions.The validity of administrative OA in British Columbia varied due to case definitions and reference standards. AOA2 is more suitable for identifying OA cases for research using this Canadian database.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.039 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.000 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".