Personal Accounts of Young-Onset Colorectal Cancer Organized as Patient-Reported Data: Protocol for a Mixed Methods Study
Bibliographic record
Abstract
BACKGROUND: Young-onset colorectal cancer is a contemporary issue in need of substantial research input. The incidence of colorectal cancer in adults younger than 50 years is rising in contrast to the decreasing incidence of this cancer in older adults. People with young-onset colorectal cancer may be at that stage of life in which they are establishing their careers, building relationships with long-term partners, raising children, and assembling a financial base for the future. A qualitative study designed to facilitate triangulation with extant quantitative patient-reported data would contribute the first comprehensive resource for understanding how this distinct patient population experiences health services and the outcomes of care throughout the patient pathway. OBJECTIVE: The aim of this study was to undertake a mixed-methods study of qualitative patient-reported data on young-onset colorectal cancer experiences and outcomes. METHODS: This is a study of web-based unsolicited patient stories recounting experiences of health services and clinical outcomes related to young-onset colorectal cancer. Personal Recollections Organized as Data (PROD) is a novel methodology for understanding patients' health experiences in order to improve care. PROD pivots qualitative data collection and analysis around the validated domains and dimensions measured in patient-reported outcome and patient-reported experience questionnaires. PROD involves 4 processes: (1) classifying attributes of the contributing patients, their disease states, their routes to diagnosis, and the clinical features of their treatment and posttreatment; (2) coding texts into the patient-reported experience and patient-reported outcome domains and dimensions, defined a priori, according to phases of the patient pathway; (3) thematic analysis of content within and across each domain; and (4) quantitative text analysis of the narrative content. RESULTS: Relevant patient stories have been identified, and permission has been obtained for use of the texts in primary research. The approval for this study was granted by the Macquarie University Human Research Ethics Committee in June 2020. The analytical framework was established in September 2020, and data collection commenced in October 2020. We will complete the analysis in March 2021 and we aim to publish the results in mid-2021. CONCLUSIONS: The findings of this study will identify areas for improvement in the PROD methodology and inform the development of a large-scale study of young-onset colorectal cancer patient narratives. We believe that this will be the first qualitative study to identify and describe the patient pathway from symptom self-identification to help-seeking through to diagnosis, treatment, and to survivorship or palliation for people with young-onset colorectal cancer. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/25056.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.080 | 0.081 |
| Meta-epidemiology (narrow) | 0.003 | 0.003 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.006 | 0.003 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.004 | 0.004 |
| Research integrity | 0.005 | 0.006 |
| Insufficient payload (model declined to judge) | 0.082 | 0.020 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".