MétaCan
Menu
Back to cohort
Record W4415221162 · doi:10.2196/80497

Online Health-Seeking Behaviors and Information Needs Among Patients With Lymphoma in China: Study of Regional and Temporal Trends

2025· article· en· W4415221162 on OpenAlexaff
Kaida Ning, Xiaoying Yang, Ling‐Li Leng, Mengting Liu, J. P. Dai, Rui Zeng, Yongshuai Hou, Rongjie Wang, Zirong Liu

Bibliographic record

VenueJournal of Medical Internet Research · 2025
Typearticle
Languageen
FieldHealth Professions
TopicHealth Literacy and Information Accessibility
Canadian institutionsUniversity of Toronto
Fundersnot available
KeywordsInformation needsGovernment (linguistics)Information seeking behaviorHealth careHealth informationNeeds assessmentPatient educationBridge (graph theory)Health informatics

Abstract

fetched live from OpenAlex

BACKGROUND: Health disparities are closely associated with socioeconomic inequalities. Although this relationship is well recognized in the context of traditional health care access, its influence on online health-seeking behaviors such as posting questions on patient forums and seeking peer responses remains poorly understood, particularly in the context of resource-limited regions. Furthermore, it is unclear what types of questions are most frequently asked online and to what extent these questions receive helpful responses. OBJECTIVE: This study aims to examine how socioeconomic status influences online health-seeking behavior by analyzing regional disparities in forum participation and their correlation with economic development. In addition, it aims to identify unmet informational needs among patients with lymphoma through large language model (LLM)-based forum thread classification and expert evaluation of forum responses by using data from the largest online blood cancer forum in China. METHODS: We analyzed over 110,000 patient-initiated forum threads posted between 2012 and 2023, covering all the provinces of mainland China. Regional trends in forum participation rates were examined and correlated with economic development, as measured by gross regional product per capita. Second, an LLM was used to classify the threads into 6 predefined topics based on their semantic content, thereby providing an overview of the topics that users cared about. Additionally, an expert manual review was conducted based on relevance, accuracy, and comprehensiveness to assess whether users' questions were adequately addressed within the forum discussions. RESULTS: Regional forum participation rates were significantly associated with levels of regional economic development (Wilcoxon rank-sum test; P<.001), with the highest participation rates in the East Coast regions. Participation rates in less-developed regions steadily increased, reflecting the growing public demand for accessible health information. LLM-based analysis revealed that most discussions centered on medical concerns such as interpreting reports and selecting treatment plans across all regions. However, only 37% (117/316) of the user questions received useful responses, underscoring persistent gaps in access to reliable information. CONCLUSIONS: To our knowledge, this study represents the most comprehensive real-world investigation to date of spontaneous online forum participation and information needs among patients with cancer. Our findings highlight the necessity for government and health care providers to implement initiatives such as artificial intelligence-driven information platforms and region-specific health education campaigns to bridge information gaps, reduce regional disparities, and improve patient outcomes across China.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.003
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.034
Threshold uncertainty score0.067

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.003
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.001
Bibliometrics0.0020.003
Science and technology studies0.0010.000
Scholarly communication0.0010.001
Open science0.0000.001
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.061
GPT teacher head0.502
Teacher spread0.441 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Medical Internet ResearchSame topicHealth Literacy and Information AccessibilityFrench-language works237,207