Bibliographic record
Abstract
Scientific inquiries into the feasibility of screening for neuroblastoma have been ongoing for 25 years. In pioneering studies, Japanese investigators showed that neuroblastomas could be detected preclinically by screening for urinary catecholamine metabolites, and that patients thus diagnosed had an almost uniformly favorable survival (1,2). Based on these results, nationwide screening for neuroblastoma was instituted in that country in the late 1980s. Currently, about one million Japanese children yearly are screened for the disease at 6 months of age (3). By the time that screening for neuroblastoma was mandated in Japan, there were many prominent voices in the Western pediatric oncology community calling for a similar “prevention” measure. However, others, concerned about several methodologic limitations of the Japanese studies, argued that neuroblastoma screening should be evaluated in a more rigorous fashion before widespread implementation (4,5). Several screening studies were instituted across Europe and North America. Two well-funded projects involved prospective population-based controlled trials to determine whether preclinical detection of neuroblastoma would affect the most important endpoint, namely mortality (6,7). These two studies, one conducted in the province of Quebec, with multiple controls throughout North America; and one conducted in Germany, with intra-German controls, have now been completed (8,9). The general concept of these trials was similar, but there were interesting logistical differences. The Quebec trial offered screening to almost 500,000 children twice, once at 3 weeks of age, to capitalize on the already in-place metabolic urinary screening program in Quebec, with greater than 90% compliance; and a second time, as a new test at 6 months of age, to mimic the Japanese experience, with a 75% compliance rate (8,10). The German study, based on preliminary results from the North American and Japanese trials (1–3,10,11), offered screening to 2,600,000 children between 9 and 18 months of age, with the hopes that preclinical detection at a later age would reduce the greatly inflated incidence of neuroblastoma that appeared to be a consequence of the earlier trials. This increased incidence was presumably from detection of neuroblastomas that normally would have regressed spontaneously without clinical detection. The Quebec trial was prospectively designed to minimize the rate of false-positives in order to spare as many young parents as possible the potential psychological burden of having their infants evaluated for cancer (10,11). Overall specificity was 99.99%; i.e., only one in 10,000 normal Quebec infants (n = 39) were screened positive and referred to one of the childhood cancer centers in that province for a neuroblastoma work-up, with negative results. The positive predictive value of the approach was 54%. Fortunately, retrospective reanalysis of urine samples from cases missed by screening confirmed that this approach did not adversely affect the sensitivity of the laboratory test. Overall sensitivity of the screening approach was 45%. The German trial, on the other hand, was designed to maximize the sensitivity of the catecholamine assays (7,9). Almost 1800 infants underwent neuroblastoma evaluation, with 149 tumors detected (positive predictive value 8%). Hence, the specificity of the approach was considerably lower than in the Quebec trial (1 in 1000 normal German infants evaluated who were false positives). However, the sensitivity was indeed elevated, to 73%. Despite these differences, the results of the Quebec and German trials were strikingly similar (8,9). In the smaller, Quebec trial, the actuarial cumulative mortality rate for neuroblastoma was 4.8 per 100,000 children at 9 years, similar to prior rates there and elsewhere. Standardized mortality ratios comparing Quebec with the various concurrent population-based control cohorts were all close to 1. In the much larger German study, screening compliance was 61%; however, nearly 1.5 million children participated. In a comparison with unscreened groups, there was a substantial amount of overdiagnosis in the screened group, more than doubling the number of cases expected. Most importantly, there was no evidence of a decrease in the incidence of late-stage neuroblastoma, that which accounts for the vast majority of neuroblastoma deaths. Although the follow-up of the German study was substantially shorter than the North American trial, the number of deaths already seen strongly suggested that there would be no usefulness to screening children, even at or after 1 year of age. The above studies nicely reemphasize a guiding principal in pediatric oncology: the importance of investigating a novel approach, be it for cancer prevention or cancer treatment, prior to its implementation. They also document that even something as simple as collecting urine from the diaper of a child can have profound implications for that child's future. In both the North American and German trials, there was a substantial overdiagnosis of neuroblastoma. Both studies noted a small number of severe adverse sequelae in the children diagnosed with neuroblastoma by screening, including the development of a “secondary” leukemia (8) and even deaths (9). One of the uncontrolled studies performed in Europe for neuroblastoma screening was done in Austria (12). That trial ended in 1999, after more than 250,000 infants had been screened. Overall in that study, 47 children were admitted to a hospital for evaluation, with 28 cases of neuroblastoma found (positive predictive value of 60%). In this issue of the Journal of Pediatric Hematology/Oncology, Dobrovoljski et al. from that trial examine the psychological burden on parents whose 19 children were referred to Austrian hospitals to rule out neuroblastoma and no such tumors were ultimately found. Parents of 16 of those children agreed to be evaluated by a semi-structured interview several years later. When questioned about the period when they were asked to bring their children to the hospital for tumor evaluation, they noted low to moderate psychological distress. However, by a median of 44 months later, 19 of the 32 parents interviewed stated that the whole experience of having a child evaluated for neuroblastoma, despite no cancer found, “had been (psychologically) striking or even very striking.” The Austrian investigators have demonstrated that parents whose children underwent screening and were found to be normal have lingering emotional distress from the experience. This is a valuable lesson for anyone contemplating other screening approaches to childhood cancer. One would like to hope that sensitive and specific screening assays could be developed in the not-too-distant future for a myriad of childhood cancers, which could potentially save thousands of lives. All such approaches would, however, need to be examined carefully: the potential for harm always exists, even when one is performing noninvasive testing. Based on the overall neuroblastoma screening results noted above, almost every instituted program in the world has abandoned this public health measure. Unfortunately, mandated screening for neuroblastoma continues in Japan. There are compelling medical, and now psychological, reasons why this practice should be strongly reconsidered:primum non nocere. Collectively we have learned much about neuroblastoma behavior and biology from the various screening trials conducted throughout the world. The results should help to minimize treatment, perhaps even to observation, in a substantial subset of infants diagnosed with early-stage neuroblastoma who have an excellent chance of tumors spontaneously maturing or regressing (13). Someday, we may have better markers of unfavorable biology neuroblastoma that can be used for preclinical detection, to favorably impact outcome. At present, however, we are nearing the final chapter in the current saga of neuroblastoma screening.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.003 | 0.004 |
| Insufficient payload (model declined to judge) | 0.041 | 0.021 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".