Mapping the Scientific Landscape of Bacterial Influence on Oral Cancer: A Bibliometric Analysis of the Last Decade’s Medical Progress
Bibliographic record
Abstract
The research domain investigating bacterial factors in the development of oral cancer from January 2013 to December 2022 was examined with a bibliometric analysis. A bibliometric analysis is a mathematical and statistical method used to examine extensive datasets. It assesses the connections between prolific authors, journals, institutions, and countries while also identifying commonly used keywords. A comprehensive search strategy identified 167 relevant articles, revealing a progressive increase in publications and citations over time. China and the United States were the leading countries in research productivity, while Harvard University and the University of Helsinki were prominent affiliations. Prolific authors such as Nezar Al-Hebshi, Tsute Chen, and Yaping Pan were identified. The analysis also highlights the contributions of different journals and identifies the top 10 most cited articles in the field, all of which focus primarily on molecular research. The article of the highest citation explored the role of a Fusobacterium nucleatum surface protein in tumor immune evasion. Other top-cited articles investigated the correlation between the oral bacteriome and cancer using 16S rRNA amplicon sequencing, showing microbial shifts associated with oral cancer development. The functional prediction analysis used by recent studies has further revealed an inflammatory bacteriome associated with carcinogenesis. Furthermore, a keyword analysis reveals four distinct research themes: cancer mechanisms, periodontitis and microbiome, inflammation and Fusobacterium, and risk factors. This analysis provides an objective assessment of the research landscape, offers valuable information, and serves as a resource for researchers to advance knowledge and collaboration in the search for the influence of bacteria on the prevention, diagnosis, and treatment of oral cancer.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.044 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.131 | 0.199 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.003 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".