Application of physician claims and hospital discharge data to the classification of carotid endarterectomy symptomatic status and to the analysis of the symptom to surgery process
Bibliographic record
Abstract
Background and purpose. The discipline of quality improvement emphasizes the use of data to improve specific and measurable attributes of performance. Using the example of carotid artery endarterectomy (CEA) for carotid artery stenosis, this thesis examined the use of hospital discharge and physician claims data for two quality improvement use cases. In Study 1, claims data were used to assemble patient cohorts based on symptomatic status at time of CEA. In Study 2, claims data were used to identify process-related causes of delayed symptomatic CEA. Methods. A single hospital’s administrative database was used to assemble a retrospective cohort of participants who had undergone CEA. Chart review data were linked with physician claims and hospital discharge data. In Study 1, a standard method of classification by hospital discharge diagnosis was compared to classification using physician claims and hospital discharge data. In Study 2, delays related to the selection of medical care activities, such as physician visits and diagnostic tests, were investigated using linear regression models of waiting time for surgery and K-means clustering for patterns of activity co-occurrence. Results. We identified 971 participants undergoing CEA at the Vancouver General Hospital (Vancouver, Canada) between January 1, 2008, and December 31, 2016. For Study 1, 729 people met inclusion/exclusion criteria and were included in diagnostic classification models (615 training, 114 test). Classification of symptomatic status using hospital discharge diagnosis codes was 32.8% (95% CI 29% – 37%) sensitive and 98.6% specific (96% – 100%). At matched 98.6% specificity, models that incorporated physician claims data were significantly more sensitive: elastic net 69.4% (59% – 82%) and random forest 78.8% (69% – 88%). In Study 2, K means clustering and linear regression analyses of process delay yielded similar results: early evaluation by emergency physicians or neurologists was associated with reduced delay, whereas carotid ultrasonography and post-imaging follow up with general practitioners or eye specialists was associated with greater delay. Conclusion. Physician claims data can be used within quality improvement projects for cohort definition and process analysis use cases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.016 | 0.078 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.008 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".