Cancer data quality and harmonization in Europe: the experience of the BENCHISTA Project – international benchmarking of childhood cancer survival by stage
Bibliographic record
Abstract
Introduction: Variation in stage at diagnosis of childhood cancers (CC) may explain differences in survival rates observed across geographical regions. The BENCHISTA project aims to understand these differences and to encourage the application of the Toronto Staging Guidelines (TG) by Population-Based Cancer Registries (PBCRs) to the most common solid paediatric cancers. Methods: PBCRs within and outside Europe were invited to participate and identify all cases of Neuroblastoma, Wilms Tumour, Medulloblastoma, Ewing Sarcoma, Rhabdomyosarcoma and Osteosarcoma diagnosed in a consecutive three-year period (2014-2017) and apply TG at diagnosis. Other non-stage prognostic factors, treatment, progression/recurrence, and cause of death information were collected as optional variables. A minimum of three-year follow-up was required. To standardise TG application by PBCRs, on-line workshops led by six tumour-specific clinical experts were held. To understand the role of data availability and quality, a survey focused on data collection/sharing processes and a quality assurance exercise were generated. To support data harmonization and query resolution a dedicated email and a question-and-answers bank were created. Results: 67 PBCRs from 28 countries participated and provided a maximally de-personalized, patient-level dataset. For 26 PBCRs, data format and ethical approval obtained by the two sponsoring institutions (UCL and INT) was sufficient for data sharing. 41 participating PBCRs required a Data Transfer Agreement (DTA) to comply with data protection regulations. Due to heterogeneity found in legal aspects, 18 months were spent on finalizing the DTA. The data collection survey was answered by 68 respondents from 63 PBCRs; 44% of them confirmed the ability to re-consult a clinician in cases where stage ascertainment was difficult/uncertain. Of the total participating PBCRs, 75% completed the staging quality assurance exercise, with a median correct answer proportion of 92% [range: 70% (rhabdomyosarcoma) to 100% (Wilms tumour)]. Conclusion: Differences in interpretation and processes required to harmonize general data protection regulations across countries were encountered causing delays in data transfer. Despite challenges, the BENCHISTA Project has established a large collaboration between PBCRs and clinicians to collect detailed and standardised TG at a population-level enhancing the understanding of the reasons for variation in overall survival rates for CC, stimulate research and improve national/regional child health plans.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".