Concordance between independent and investigator assessment of disease-free survival (DFS) in the APACT trial.
Bibliographic record
Abstract
4618 Background: APACT was a phase III trial of adjuvant nab-paclitaxel + gemcitabine ( nab-P + Gem) vs Gem alone in patients with resected pancreatic cancer (PC) and the first adjuvant PC trial to use independently assessed DFS as the primary endpoint (DFS by investigator review was a prespecified sensitivity analysis). We examined concordance between independent and investigator DFS review. Methods: For the independent assessment, reviewers determined recurrence by computed tomography or magnetic resonance imaging but were blinded to treatment and clinical data. Investigator-assessed DFS was based on all available data. Concordance was summarized by κ statistics. Patients who did not have recurrence or were alive were censored at the last tumor assessment date with disease-free status or the randomization date if the last tumor assessment with disease-free status was missing. Patients who received new anticancer therapy or cancer-related surgery prior to recurrence or death were censored at the date of last tumor assessment with disease-free status prior to the start of new anticancer therapy or cancer-related surgery or the randomization date if the last tumor assessment date with disease-free status prior to the start of subsequent new anticancer therapy or cancer-related surgery was missing. All censoring rules were the same for analysis of DFS by independent and investigator review. Results: Median DFS by independent review was 19.4 ( nab-P + Gem) vs 18.8 (Gem) months (hazard ratio [HR] 0.88; 95% CI, 0.73 - 1.06; P = 0.18); median investigator-assessed DFS was 16.6 ( nab-P + Gem) vs 13.7 (Gem) months (HR 0.82; 95% CI, 0.69 - 0.97; nominal P = 0.017). Moderate concordance was found between independent- and investigator-assessed DFS (Table); similar results were observed in the nab-P + Gem (concordance, 78%; κ coefficient, 0.56) and Gem alone (concordance, 76%; κ coefficient, 0.53) arms. Conclusions: The results reflect the complexities of defining the recurrence timepoint accurately and suggest that radiological review in the absence of clinical context is suboptimal for recurrence detection in resected PC. These findings may inform future clinical trial design. Registration: EudraCT (2013-003398-91); ClinicalTrials.gov (NCT01964430). Clinical trial information: NCT01964430 . [Table: see text]
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.193 | 0.267 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".