Evaluation of A Clinical Decision Support Tool for Selecting Optimal Rehabilitation Intervention for Injured Workers
Bibliographic record
Abstract
Objective: To evaluate the concurrent validity of a newly developed clinical decision support tool (Work Assessment Triage Tool, WATT) by comparing the rehabilitation interventions determined using the WATT with the current gold standard–clinician recommendations. Methods: This is a secondary data analysis study. Data were collected in a clinical trial conducted previously at the Workers’ Compensation Board of Alberta rehabilitation facility. A variety of statistical methods were used to compare recommendations for rehabilitation strategies determined using the WATT, clinician recommendations, actual programs claimants undertook and return-to-work outcomes. Analyses included percent agreement, crosstabs, and likelihood ratios. Results: Percent agreement between clinician recommendations and WATT recommendations were low (r = 0.19) to moderate (r = 0.46). The WATT does not appear to improve upon clinician recommendations as only half of the RTW claimants whose actual rehabilitation programs did not match those of the clinician recommendations, matched recommendations identified using WATT. Discussions: Contrary to internal validation demonstrating that the WATT outperformed clinician recommendations; results of the external validation of the WATT were not as promising. Findings do not provide evidence of concurrent validity of the WATT against the current gold standard. Four possible reasons could explain the results: (1) important differences were observed in claimant characteristics between the original WATT development data and our validation dataset; (2) insufficient data for claimants who failed RTW and those with successful RTW whose actual rehabilitation program did not match with the clinician recommendations; (3) data processing techniques that were used to overcome rehabilitation class imbalance when building the WATT, which may contribute to errors in the WATT recommendations; (4) clinician recommendations conflicted somewhat with existing evidence as some rehabilitation programs that were highly supported by research evidence (i.e. workplace interventions) were rarely recommended by clinicians in our validation dataset. Conclusion: WATT recommendations do not concur with clinician recommendations. With respect to concurrent validity, no conclusion can be drawn as to which method, WATT or clinician judgment, provides better recommendations for return-to-work in actual practice. Further research is needed to resolve this uncertainty.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.062 | 0.238 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.004 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".