Accuracy of Self-Reported Breast Cancer Information among Women from the Ontario Site of the Breast Cancer Family Registry
Bibliographic record
Abstract
Obtaining complete medical record information can be challenging and expensive in breast cancer studies. The current literature is limited with respect to the accuracy of self-report and factors that may influence this. We assessed the agreement between self-reported and medical record breast cancer information among women from the Ontario site of the Breast Cancer Family Registry. Women aged 20-69 years diagnosed with incident breast cancer 1996-1998 were identified from the Ontario Cancer Registry, sampled on age and family history. We calculated kappa statistics, proportion correct, sensitivity, specificity, and positive and negative predictive values and conducted unconditional logistic regression to examine whether characteristics of the women influenced agreement. The proportions of women who correctly reported having received a broad category of therapy (hormone therapy, chemotherapy, radiation, or surgery) as well as sensitivity and specificity were above 90%, and the kappa statistics were above 0.80. The specific type of hormonal or chemotherapy was reported with low-to-moderate agreement. Aside from recurrence, no factors were consistently associated with agreement. Thus, most women were able to accurately report broad categories of treatment but not necessarily specific treatment types. The finding of this study can aid researchers in the use and design of self-administered treatment questionnaires.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".