Abstract 4513: St. Jude Survivorship Portal: A data portal for storing, analyzing, and sharing large and complex cancer survivorship datasets
Bibliographic record
Abstract
Abstract Survivors of childhood cancer are at risk for developing various adverse health conditions as adults that are attributable to the cancer and treatments they were exposed to as children. Cancer survivorship research relies on large-scale, longitudinal studies that generate a wide range of demographic, clinical, and genetic data on cancer survivors at multiple time points. To maximize the utility of these comprehensive datasets, we must be able to store and share these datasets in a web-based environment that can be accessed by the broader survivorship research community. Furthermore, this environment should be integrated with analytical tools for performing statistical analyses on the stored data without needing to download the data and import it into third-party analytical software. To address this need, we have created the St. Jude Survivorship Portal (https://survivorship.stjude.cloud > Clinical Data Browser), a web-based data portal for exploring, sharing, and analyzing data from survivors of pediatric cancer. The portal hosts data from two large cohorts of pediatric cancer survivors: the St. Jude Lifetime Cohort Study and the Childhood Cancer Survivor Study. The data stored on the portal consists of demographic data, clinical data, including cancer diagnosis, cancer treatment, clinical outcomes, and patient-reported data, and genetic data, including whole-genome-sequencing-derived genotypes and published polygenic risk scores computed for >500 traits. This data is organized hierarchically in a data dictionary that can be easily explored by the user. Charts and plots of variables can be quickly created, customized, and stratified with other variables, all within the portal environment. Statistical analyses, including cumulative incidence analysis and regression analysis, may also be performed within the portal. In cumulative incidence analysis, users can analyze the incidence of a variety of CTCAE-graded adverse events (e.g., cardiovascular dysfunction, neurological disorders, subsequent neoplasms) in survivors and can also compare them across different survivor populations defined by other variables. In regression analysis, users have the option to perform either a linear, logistic, or cox regression analysis and may use any of the demographic, clinical, or genetic variables on the portal as outcome or explanatory variables in the analysis. In this way, users can assess any risk factor associations within a survivor cohort and generate predictive models for outcomes of interest. Lastly, we also provide the user with the option to download the data on the portal for use in any future analyses. The St. Jude Survivorship Portal provides a comprehensive, powerful, and easy-to-use interface for sharing and analyzing childhood cancer survivorship data that will serve as a valuable research tool for the broader survivorship research community. Citation Format: Gavriel Matt, Edgar Sioson, Jian Wang, Congyu Lu, Airen Zaldivar Peraza, Karishma Gangwani, Alex Acic, Jaimin Patel, Robin Paul, Colleen Reilly, Kyla Shelton, Qi Liu, Weiyu Qiu, Cindy Im, Zhaoming Wang, Carmen L. Wilson, Nickhill Bhakta, Kirsten Ness, Gregory T. Armstrong, Melissa M. Hudson, Leslie L. Robison, Jinghui Zhang, Yutaka Yasui, Xin Zhou. St. Jude Survivorship Portal: A data portal for storing, analyzing, and sharing large and complex cancer survivorship datasets. [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 1 (Regular and Invited Abstracts); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(7_Suppl):Abstract nr 4513.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.030 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.007 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.006 |
| Open science | 0.004 | 0.007 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.127 | 0.083 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".