Creation of a Diagnostic Wait Times Measurement Framework Based on Evidence and Consensus
Bibliographic record
Abstract
PURPOSE: Public reporting of wait times worldwide has to date focused largely on treatment wait times and is limited in its ability to capture earlier parts of the patient journey. The interval between suspicion and diagnosis or ruling out of cancer is a complex phase of the cancer journey. Diagnostic delays and inefficient use of diagnostic imaging procedures can result in poor patient outcomes, both physical and psychosocial. This study was designed to develop a framework that could be adopted for multiple disease sites across different jurisdictions to enable the measurement of diagnostic wait times and diagnostic delay. METHODS: Diagnostic benchmarks and targets in cancer systems were explored through a targeted literature review and jurisdictional scan. Cancer system leaders and clinicians were interviewed to validate the information found in the jurisdictional scan. An expert panel was assembled to review and, through a modified Delphi consensus process, provide feedback on a diagnostic wait times framework. RESULTS: The consensus process resulted in agreement on a measurement framework that identified suspicion, referral, diagnosis, and treatment as the main time points for measuring this critical phase of the patient journey. CONCLUSIONS: This work will help guide initiatives designed to improve patient access to health services by developing an evidence-based approach to standardization of the various waypoints during the diagnostic pathway. The diagnostic wait times measurement framework provides a yardstick to measure the performance of programs that are designed to manage and expedite care processes between referral and diagnosis or ruling out of cancer.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.122 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".