“Googling” for Cancer: An Infodemiological Assessment of Online Search Interests in Australia, Canada, New Zealand, the United Kingdom, and the United States
Bibliographic record
Abstract
BACKGROUND: The infodemiological analysis of queries from search engines to shed light on the status of various noncommunicable diseases has gained increasing popularity in recent years. OBJECTIVE: The aim of the study was to determine the international perspective on the distribution of information seeking in Google regarding "cancer" in major English-speaking countries. METHODS: We used Google Trends service to assess people's interest in searching about "Cancer" classified as "Disease," from January 2004 to December 2015 in Australia, Canada, New Zealand, the United Kingdom, and the United States. Then, we evaluated top cities and their relative search volumes (SVs) and country-specific "Top searches" and "Rising searches." We also evaluated the cross-country correlations of SVs for cancer, as well as rank correlations of SVs from 2010 to 2014 with the incidence of cancer in 2012 in the abovementioned countries. RESULTS: From 2004 to 2015, the United States (relative SV [from 100]: 63), Canada (62), and Australia (61) were the top countries searching for cancer in Google, followed by New Zealand (54) and the United Kingdom (48). There was a consistent seasonality pattern in searching for cancer in the United States, Canada, Australia, and New Zealand. Baltimore (United States), St John's (Canada), Sydney (Australia), Otaika (New Zealand), and Saint Albans (United Kingdom) had the highest search interest in their corresponding countries. "Breast cancer" was the cancer entity that consistently appeared high in the list of top searches in all 5 countries. The "Rising searches" were "pancreatic cancer" in Canada and "ovarian cancer" in New Zealand. Cross-correlation of SVs was strong between the United States, Canada, and Australia (>.70, P<.01). CONCLUSIONS: Cancer maintained its popularity as a search term for people in the United States, Canada, and Australia, comparably higher than New Zealand and the United Kingdom. The increased interest in searching for keywords related to cancer shows the possible effectiveness of awareness campaigns in increasing societal demand for health information on the Web, to be met in community-wide communication or awareness interventions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.008 | 0.014 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".