The Validity of Administrative Data To Classify Patients with Spinal Column and Cord Injuries
Bibliographic record
Abstract
International Classification of Diseases (ICD) codes are used to document patient morbidity in administrative databases. Although administrative data are used for research purposes, the validity of the data to accurately describe clinical diagnostic information is uncertain. We compared the clinical diagnoses for spinal cord and column injuries from a longitudinal patient registry, the Rick Hansen Spinal Cord Injury Registry (RHSCIR), to the ICD-10 spinal injury codes from the Discharge Abstract Database (DAD) at one institution. There were 603 RHSCIR participants with data describing the spinal cord injury, and 341 had data on the spinal column injury. The validity of DAD data to describe spinal injuries was evaluated using the sensitivity and positive predictive values of specific ICD-10 codes; 5.3% of the spinal column injuries and 10.9% of the spinal cord injuries documented in RHSCIR were missed in data from the DAD using ICD-10 codes. The most problematic spinal column ICD-10 code was the dislocation of the cervical vertebra (S13.1); only 14.0% of the dislocations of the cervical vertebrae in RHSCIR were correctly coded in the DAD. The most problematic spinal cord injury ICD-10 code was the incomplete lesion of the lumbar spinal cord (S34.1X); 66.7% of incomplete lesions of the lumbar spinal cord in RHSCIR were correctly coded in the DAD. The validity of DAD data to code spinal injuries is variable, and cannot be reliably used to classify all types of spinal injuries. Patient registries, such as RHSCIR, should be used if accurate detailed diagnostic data are required.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".