Development of Quality Metrics to Evaluate Pediatric Hematologic Oncology Care in the Outpatient Setting
Bibliographic record
Abstract
Abstract Introduction: Systems to quantify and incentivize quality of care (QoC) have been developed in multiple healthcare settings. In pediatric oncology, lists of QoC metrics or recommendations have been procured through consensus methodologies such as the Delphi process. To date, no QoC metrics have been developed for outpatient pediatric oncology. Objectives: The aim of this study was to develop a list of QoC metrics for the leukeumia-lymphoma (LL) clinic at the Hospital for Sick Children in Toronto, using a consensus process that could be adapted to other clinic settings. Methods: A modified Delphi process following the American Society of Clinical Oncology (ASCO) guidelines was used to generate consensus on a list of QoC metrics (Loblaw et al., 2012). A Medline-Ovid search was conducted for quality indicators, metrics and recommendations relevant to pediatric oncology. Results were screened for (a) system-level metrics that could be translated to a clinic level and (b) clinic-level recommendations that could be converted to measurable quantities. Additional metrics outside the literature search were considered. A provisional list was compiled and circulated electronically to local stakeholders, including medical and nursing staff (n=10). Stakeholders ranked each metric on a 5-point Likert scale based on importance and feasibility of measurement (round 1). Stakeholders provided feedback on the metrics and suggested additional metrics. Median, interquartile range and full ranges were calculated for each metric. A metric was considered to reach consensus if the percent of respondents ranking within two consecutive scores was ≥70%. Results and comments from round 1 were re-circulated to stakeholders in personalized reports. This allowed each stakeholder to compare his or her previous scores with overall scores for each metric. Stakeholders were asked to re-rank each metric (round 2). Results: The literature search yielded 2 relevant publications from which a provisional list of 27 metrics was generated. Metrics were grouped into 7 categories (Table 1). In round 1, 19/27 (70%) metrics reached consensus. Stakeholders’ comments resulted in 4 new metrics and edits to 8 original metrics. All metrics were included in round 2 for a total of 31. Twenty-four of 31 (77%) metrics reached consensus after round 2 (Table 1). Thirteen were chosen for the final list based on highest consensus scores, highest interquartile and full ranges, and minimizing redundancy. Conclusion: This study demonstrates the feasibility of using a modified Delphi process to generate QoC metrics for a pediatric hematology oncology clinic, and provides a model other clinics may employ for local use. The final metrics will be used to evaluate the quality of care in the LL clinic, and to identify areas for improvement in clinic function. Figure 1 Figure 1. Disclosures No relevant conflicts of interest to declare.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".