Cancer data quality and harmonization in Europe: the experience of the BENCHISTA Project – international benchmarking of childhood cancer survival by stage
Bibliographic record
Abstract
Introduction: Variation in stage at diagnosis of childhood cancers (CC) may explain differences in survival rates observed across geographical regions. The BENCHISTA project aims to understand these differences and to encourage the application of the Toronto Staging Guidelines (TG) by Population-Based Cancer Registries (PBCRs) to the most common solid paediatric cancers. Methods: PBCRs within and outside Europe were invited to participate and identify all cases of Neuroblastoma, Wilms Tumour, Medulloblastoma, Ewing Sarcoma, Rhabdomyosarcoma and Osteosarcoma diagnosed in a consecutive three-year period (2014-2017) and apply TG at diagnosis. Other non-stage prognostic factors, treatment, progression/recurrence, and cause of death information were collected as optional variables. A minimum of three-year follow-up was required. To standardise TG application by PBCRs, on-line workshops led by six tumour-specific clinical experts were held. To understand the role of data availability and quality, a survey focused on data collection/sharing processes and a quality assurance exercise were generated. To support data harmonization and query resolution a dedicated email and a question-and-answers bank were created. Results: 67 PBCRs from 28 countries participated and provided a maximally de-personalized, patient-level dataset. For 26 PBCRs, data format and ethical approval obtained by the two sponsoring institutions (UCL and INT) was sufficient for data sharing. 41 participating PBCRs required a Data Transfer Agreement (DTA) to comply with data protection regulations. Due to heterogeneity found in legal aspects, 18 months were spent on finalizing the DTA. The data collection survey was answered by 68 respondents from 63 PBCRs; 44% of them confirmed the ability to re-consult a clinician in cases where stage ascertainment was difficult/uncertain. Of the total participating PBCRs, 75% completed the staging quality assurance exercise, with a median correct answer proportion of 92% [range: 70% (rhabdomyosarcoma) to 100% (Wilms tumour)]. Conclusion: Differences in interpretation and processes required to harmonize general data protection regulations across countries were encountered causing delays in data transfer. Despite challenges, the BENCHISTA Project has established a large collaboration between PBCRs and clinicians to collect detailed and standardised TG at a population-level enhancing the understanding of the reasons for variation in overall survival rates for CC, stimulate research and improve national/regional child health plans.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.158 | 0.096 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.010 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.005 | 0.003 |
| Open science | 0.004 | 0.009 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".