Association Versus Causation Versus Quality Improvement: Setting Benchmarks for Lymph Node Evaluation in Colon Cancer
Bibliographic record
Abstract
There has been substantial attention and interest directed toward improving the quality of medical care in the United States; the need for quality improvement has reached the consideration of policy makers, providers, payers, and patients. In response to congressional mandates, the Institute of Medicine launched the Redesigning Health Insurance Performance Measures, Payment, and Performance Improvement Project ( 1 ), with the goal of accelerating the diffusion and pace of quality improvement efforts. Specific policies have been promoted to improve care, including measurement and reporting of performance data, payment incentives, and quality improvement initiatives. Measures in oncology are under active development, and as this process evolves, it is likely that implementation of performance measures will become mandatory and that the scope will broaden. Lymph node evaluation is a frequently discussed potential quality measure for colon cancer, and benchmarks for adequacy of lymph node evaluation have been proposed. As Chang et al. ( 2 ) point out, “the number of lymph nodes recovered from a patient with colon cancer has been identified as a potentially important measure of the quality of cancer care by many organizations, including the American College of Surgeons, the American Society of Clinical Oncology, the National Comprehensive Cancer Network, the National Quality Forum, healthcare insurance providers, and others.” This paper, a systematic review of the evidence associating lymph node harvest in colon cancer and clinical outcomes is, therefore, both timely and topical. In a pooled analysis including more than 60 000 patients ( 2 ), the authors found that 16 of 17 national and international studies demonstrated improved survival as the number of lymph nodes evaluated increased in patients with stage II colon cancer. In addition, four of six studies reported a positive association between lymph node number and survival among patients with stage III colon cancer. The authors conclude that given the evidence, lymph node evaluation deserves consideration as a quality measure for colon cancer care. However, before lymph node benchmarks are established as a quality measure, two important questions must be addressed. First, who or what is being evaluated when we report lymph node counts — the surgeon, the pathologist, the hospital, the patient, or even the tumor?
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".