Policy Directives in Cancer, Organizational Responses, and the Need for Evaluation
Bibliographic record
Abstract
In 2007, the US Department of Veterans Affairs (VA) issued a directive encouraging colorectal cancer (CRC) screening for veterans who were deemed to be at average or high risk for CRC. In response, in the year 2008, administrators and clinical leaders in Veterans Integrated Service Network 7 implemented a computerized clinical reminder system known as Oncology Watch (OncWatch) to increase CRC screening rates and to improve the use of CRC colonoscopy diagnostic and surveillance services. The system was implemented in all eight of the Network 7 hospitals, whereas no similar interventions were initiated at the remaining 121 VA sites. In the article that accompanies this editorial, Bian et al report on an evaluation of the organizational decision to initiate OncWatch in Network 7 sites. They conclude that OncWatch had no impact on CRC screening rates in Network 7, and that OncWatch may have unintentionally diverted “limited VA colonoscopy capacity from average-risk screening to higher-risk screening and to CRC surveillance.” This editorial will comment on the latter conclusion, but will first highlight the importance of this evaluation of an organizational response to a policy directive. In the article by Bian et al, intervention cohorts (patients using services at any Network 7 hospital or affiliated clinic) and control cohorts (patients using services at any of the remaining 121 VA sites) were created for each of the years 2006 and 2007 (preintervention period) and 2009 and 2010 (postintervention period). For reasons well supported in the methods, only veterans age 50 to 64 years with average risk for CRC were included in the cohorts. The authors defined screening as fecal occult blood testing in the year under review; flexible sigmoidoscopy or barium enema in the year under review or during the 5 years before; or colonoscopy in the year under review or during the 9 years before. The article includes four key observations. First, screening rates were low for all years; the highest rate was only 37.6% in the intervention cohort in the year 2006. Second, use of OncWatch among Network 7 sites was associated with a 2.2% drop in the likelihood of overall CRC screening adherence. Third, use of OncWatch was associated with a 5.6% drop in the likelihood of screening with colonoscopy, and this absolute 5.6% drop represented a relative 26.1% drop in the use of colonoscopy as a screening modality. Fourth, use of OncWatch was associated with an overall increase of 3.6 colonoscopies per 100 veterans, suggesting that the drop in average-risk screening colonoscopy was more than compensated for by an increased use of colonoscopy for high-risk screening, diagnostics, or surveillance among veterans cared for in Network 7 hospitals. In any health care system, well-intentioned policy makers, administrators, and clinical leaders will develop and mandate directives that are designed to close a measured or perceived quality gap. Depending on the context of the higher-level organization, directives may take the form of official policy, a document to encourage change, or a general statement. Closer to the frontlines and again dependent on context, suborganizations or individuals may implement practice changes in response to such directives. Unfortunately, this is often where the process ends. An expert evaluation of the impact of organizational changes on patient care or relevant quality markers is the exception and not the rule. Considering the inevitable absolute dollars and opportunity costs involved in selecting specific aspects of disease management for attention and action, it is recommended that such evaluations become the rule for the benefit of sponsoring policy agencies, sub-actors implementing organizational changes, and, ultimately, patient care. Volume-outcome studies and resulting policy directives provide an example of the importance of subsequent evaluation. In response to findings of positive volume-outcome relationships in major cancer surgery procedures (ie, superior patient outcomes are associated with higher v lower hospital procedure volumes), policy-making groups around the world have encouraged or mandated the regionalization of an assortment of surgical procedures to high-volume centers. But the few evaluations of such regionalization efforts provide surprising results. In a study using data from the Canadian provinces of Ontario and Quebec, Simunovic et al recently found that regionalization trends in pancreas cancer surgery were likely not influenced by an Ontario policy directive from the provincial cancer agency, and that regionalization did not guarantee improved patient outcomes. Similarly, in Washington state, regionalization of various major cancer procedures was not associated with improved patient outcomes. Given the results of these evaluations, policy agents should not rely on regionalization alone as a panacea for improved patient care in cancer surgery. With respect to the article by Bian et al, the VA mandated CRC screening for averageand high-risk veterans. Stakeholders in Network 7 responded by implementing OncWatch. Fortunately for all involved, Bian et al have provided a high-quality evaluation of this Network 7 decision. Although it is disappointing that OncWatch did JOURNAL OF CLINICAL ONCOLOGY E D I T O R I A L VOLUME 30 NUMBER 32 NOVEMBER 1
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.404 | 0.525 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.004 | 0.007 |
| Science and technology studies | 0.012 | 0.035 |
| Scholarly communication | 0.028 | 0.039 |
| Open science | 0.008 | 0.012 |
| Research integrity | 0.018 | 0.029 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".