Development of the Czech Childhood Cancer Information System: Data Analysis and Interactive Visualization
Bibliographic record
Abstract
BACKGROUND: The knowledge of cancer burden in the population, its time trends, and the possibility of international comparison is an important starting point for cancer programs. A reliable interactive tool describing cancer epidemiology in children and adolescents has been nonexistent in the Czech Republic until recently. OBJECTIVE: The goal of this study is to develop a new web portal entitled the Czech Childhood Cancer Information System (CCCIS), which would provide information on childhood cancer epidemiology in the Czech Republic. METHODS: Data on childhood cancers have been obtained from the Czech National Cancer Registry. These data were validated using the clinical database of childhood cancer patients and subsequently combined with data from the National Register of Hospitalised Patients and with data from death certificates. These validated data were then used to determine the incidence and survival rates of childhood cancer patients aged 0 to 19 years who were diagnosed in the period 1994 to 2016 (N=9435). Data from death certificates were used to monitor long-term mortality trends. The technical solution is based on the robust PHP development Symfony framework, with the PostgreSQL system used to accommodate the data basis. RESULTS: The web portal has been available for anyone since November 2019, providing basic information for experts (ie, analyses and publications) on individual diagnostic groups of childhood cancers. It involves an interactive tool for analytical reporting, which provides information on the following basic topics in the form of graphs or tables: incidence, mortality, and overall survival. Feedback was obtained and the accuracy of outputs published on the CCCIS portal was verified using the following methods: the validation of the theoretical background and the user testing. CONCLUSIONS: We developed software capable of processing data from multiple sources, which is freely available to all users and makes it possible to carry out automated analyses even for users without mathematical background; a simple selection of a topic to be analyzed is required from the user.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.024 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.006 | 0.005 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.005 | 0.003 |
| Open science | 0.002 | 0.005 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.032 | 0.015 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".