Mapping intellectual structure and research hotspots of cancer studies in primary health care: A machine-learning-based analysis
Bibliographic record
Abstract
In the contemporary fight against cancer, primary health care (PHC) services hold a significant and critical position within the healthcare system. This study, as one of the most detailed investigations into cancer research in primary care, comprehensively evaluates cancer studies from the perspective of PHC using bibliometric techniques and machine learning. The dataset for the analyses was sourced from the Web of Science (WoS) Core Collection database on March 20, 2024. The Bibliometrix package within the R programming environment, alongside the Biblioshiny application, and VOSViewer software were employed for the bibliometric analyses. In this study, Latent Dirichlet Allocation was utilized as a prominent topic modeling algorithm. The implementation of this technique utilized Python along with the SciKit-Learn and Gensim libraries, ensuring robust model development and evaluation. The 2040 articles were produced by a total of 6705 different authors, 2166 different affiliations, and 75 different countries. Cancer survivors are more vulnerable and need more sensitive health services. The most intensively studied 3 cancer types in the PHC, listed by prevalence, are colorectal cancer, breast cancer, and cervical cancer. Additionally, prominent research topics in PHC include cancer screening, diagnosis, early detection, prevention, education, genetic factors and family history, risk factors, symptoms/signs, preventive medicine, referral and consultation, chronic disease management and health services research for cancer patients, health care disparities, palliative care, and communication with patients in PHC. Family physicians, being the first point of contact with the public, play a crucial role in preventing cancer cases, caring for patients with active cancer diagnoses, supporting cancer survivors in their post-cancer lives, and identifying and referring cancer cases at the earliest stages. However, cancer has many types, each with its own distinct symptoms, as well as similar types to each other. At this point, periodic educational training for doctors on cancer by health authorities, regular publication of cancer-related guidance resources by the central healthcare system, development of integrated decision support tools used by physicians during patient care, and the creation of informative mobile applications for cancer prevention or post-cancer life for patients have been considered highly critical.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.037 | 0.197 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.004 |
| Bibliometrics | 0.088 | 0.133 |
| Science and technology studies | 0.003 | 0.003 |
| Scholarly communication | 0.008 | 0.005 |
| Open science | 0.002 | 0.008 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".