Global colorectal cancer research, 2007‐2021: Outputs and funding
Bibliographic record
Abstract
The purpose of this study was to provide an evidence base for colorectal cancer research activity that might influence policy, mainly at the national level. Improvements in healthcare delivery have lengthened life expectancy, but within a situation of increased cancer incidence. The disease burden of CRC has risen significantly, particularly in Africa, Asia and Latin America. Research is key to its control and reduction, but few studies have delineated the volume and funding of global research on CRC. We identified research papers in the Web of Science (WoS) from 2007 to 2021, and determined the contributions of the leading countries, the research domains studied, and their sources of funding. We identified 62 716 papers, representing 5.7% of all cancer papers. This percentage was somewhat disproportionate to the disease burden (7.7% in 2015), especially in Eastern Europe. International collaboration increased over the time period in almost all countries except in China. Genetics, surgery and prognosis were the leading research domains. However, research on palliative care and quality-of-life in CRC was lacking. In Western Europe, the main funding source was the charity sector, particularly in the UK, but in most other countries government played the leading role, especially in China and the USA. There was little support from industry. Several Asian countries provided minimal contestable funding, which may have reduced the impact of their CRC research. Certain countries must perform more CRC research overall, especially in domains such as screening, palliative care and quality-of-life. The private-non-profit sector should be an alternative source of support.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.066 | 0.171 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.031 | 0.048 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.011 | 0.006 |
| Open science | 0.002 | 0.007 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.029 | 0.019 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".