Web-Based Information Seeking Behaviors of Low-Literacy Hispanic Survivors of Breast Cancer: Observational Pilot Study
Bibliographic record
Abstract
BACKGROUND: Internet searching is a useful tool for seeking health information and one that can benefit low-literacy populations. However, low-literacy Hispanic survivors of breast cancer do not normally search for health information on the web. For them, the process of searching can be frustrating, as frequent mistakes while typing can result in misleading search results lists. Searches using voice (dictation) are preferred by this population; however, even if an appropriate result list is displayed, low-literacy Hispanic women may be challenged in their ability to fully understand any individual article from that list because of the complexity of the writing. OBJECTIVE: This observational study aims to explore and describe web-based search behaviors of Hispanic survivors of breast cancer by themselves and with their caregivers, as well as to describe the challenges they face when processing health information on the web. METHODS: We recruited 7 Hispanic female survivors of breast cancer. They had the option to bring a caregiver. Of the 7 women, 3 (43%) did, totaling 10 women. We administered the Health LiTT health literacy test, a demographic survey, and a breast cancer knowledge assessment. Next, we trained the participants to search on the web with either a keyboard or via voice. Then, they had to find information about 3 guided queries and 1 free-form query related to breast cancer. Participants were allowed to search in English or in Spanish. We video and audio recorded the computer activity of all participants and analyzed it. RESULTS: We found web articles to be written for a grade level of 11.33 in English and 7.15 in Spanish. We also found that most participants preferred searching using voice but struggled with this modality. Pausing while searching via voice resulted in incomplete search queries, as it confused the search engine. At other times, background noises were detected and included in the search. We also found that participants formulated overly general queries to broaden the results list hoping to find more specific information. In addition, several participants considered their queries satisfied based on information from the snippets on the result lists alone. Finally, participants who spent more time reviewing articles scored higher on the health literacy test. CONCLUSIONS: Despite the problems of searching using speech, we found a preference for this modality, which suggests a need to avoid potential errors that could appear in written queries. We also found the use of general questions to increase the chances of answers to more specific concerns. Understanding search behaviors and information evaluation strategies for low-literacy Hispanic women survivors of breast cancer is fundamental to designing useful search interfaces that yield relevant and reliable information on the web.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".