MétaCan
Menu
← Back to cohort
Record W4311609625 · doi:10.2196/36755

Sample Bias in Web-Based Patient-Generated Health Data of Dutch Patients With Gastrointestinal Stromal Tumor: Survey Study

2022· article· en· W4311609625 on OpenAlexvenueno aff
Anne Dirkson, Dide den Hollander, Suzan Verberne, Ingrid M.E. Desar, Olga Husson, Winette T.A. van der Graaf, Astrid W. Oosten, Anna K.L. Reyners, Neeltje Steeghs, Wouter van Loon, G. van Oortmerssen, Hans Gelderblom, Wessel Kraaij

Bibliographic record

VenueJMIR Formative Research · 2022
Typearticle
Languageen
FieldSocial Sciences
TopicSocial Media in Health Education
Canadian institutionsnot available
FundersEuropean Organisation for Research and Treatment of Cancer
KeywordsGiSTMedicinePopulationSocial mediaSocioeconomic statusFamily medicineRepresentativeness heuristicLogistic regressionInternal medicineDemographyPsychologyStromal cellEnvironmental healthWorld Wide Web

Abstract

fetched live from OpenAlex

BACKGROUND: Increasingly, social media is being recognized as a potential resource for patient-generated health data, for example, for pharmacovigilance. Although the representativeness of the web-based patient population is often noted as a concern, studies in this field are limited. OBJECTIVE: This study aimed to investigate the sample bias of patient-centered social media in Dutch patients with gastrointestinal stromal tumor (GIST). METHODS: A population-based survey was conducted in the Netherlands among 328 patients with GIST diagnosed 2-13 years ago to investigate their digital communication use with fellow patients. A logistic regression analysis was used to analyze clinical and demographic differences between forum users and nonusers. RESULTS: Overall, 17.9% (59/328) of survey respondents reported having contact with fellow patients via social media. Moreover, 78% (46/59) of forum users made use of GIST patient forums. We found no statistically significant differences for age, sex, socioeconomic status, and time since diagnosis between forum users (n=46) and nonusers (n=273). Patient forum users did differ significantly in (self-reported) treatment phase from nonusers (P=.001). Of the 46 forum users, only 2 (4%) were cured and not being monitored; 3 (7%) were on adjuvant, curative treatment; 19 (41%) were being monitored after adjuvant treatment; and 22 (48%) were on palliative treatment. In contrast, of the 273 patients who did not use disease-specific forums to communicate with fellow patients, 56 (20.5%) were cured and not being monitored, 31 (11.3%) were on curative treatment, 139 (50.9%) were being monitored after treatment, and 42 (15.3%) were on palliative treatment. The odds of being on a patient forum were 2.8 times as high for a patient who is being monitored compared with a patient that is considered cured. The odds of being on a patient forum were 1.9 times as high for patients who were on curative (adjuvant) treatment and 10 times as high for patients who were in the palliative phase compared with patients who were considered cured. Forum users also reported a lower level of social functioning (84.8 out of 100) than nonusers (93.8 out of 100; P=.008). CONCLUSIONS: Forum users showed no particular bias on the most important demographic variables of age, sex, socioeconomic status, and time since diagnosis. This may reflect the narrowing digital divide. Overrepresentation and underrepresentation of patients with GIST in different treatment phases on social media should be taken into account when sourcing patient forums for patient-generated health data. A further investigation of the sample bias in other web-based patient populations is warranted.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.017
metaresearch head score (Gemma)0.054
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.983
Threshold uncertainty score0.091

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0170.054
Meta-epidemiology (narrow)0.0000.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.003
Science and technology studies0.0010.001
Scholarly communication0.0010.001
Open science0.0010.002
Research integrity0.0010.000
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.411
GPT teacher head0.515
Teacher spread0.104 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designObservational
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Formative Research→Same topicSocial Media in Health Education→French-language works237,207→