Worldwide inequality in access to full text scientific articles: the example of ophthalmology
Bibliographic record
Abstract
BACKGROUND: The problem of access to medical information, particularly in low-income countries, has been under discussion for many years. Although a number of developments have occurred in the last decade (e.g., the open access (OA) movement and the website Sci-Hub), everyone agrees that these difficulties still persist very widely, mainly due to the fact that paywalls still limit access to approximately 75% of scholarly documents. In this study, we compare the accessibility of recent full text articles in the field of ophthalmology in 27 established institutions located worldwide. METHODS: A total of 200 references from articles were retrieved using the PubMed database. Each article was individually checked for OA. Full texts of non-OA (i.e., "paywalled articles") were examined to determine whether they were available using institutional and Hinari access in each institution studied, using "alternative ways" (i.e., PubMed Central, ResearchGate, Google Scholar, and Online Reprint Request), and using the website Sci-Hub. RESULTS: The number of full texts of "paywalled articles" available using institutional and Hinari access showed strong heterogeneity, scattered between 0% full texts to 94.8% (mean = 46.8%; SD = 31.5; median = 51.3%). We found that complementary use of "alternative ways" and Sci-Hub leads to 95.5% of full text "paywalled articles," and also divides by 14 the average extra costs needed to obtain all full texts on publishers' websites using pay-per-view. CONCLUSIONS: The scant number of available full text "paywalled articles" in most institutions studied encourages researchers in the field of ophthalmology to use Sci-Hub to search for scientific information. The scientific community and decision-makers must unite and strengthen their efforts to find solutions to improve access to scientific literature worldwide and avoid an implosion of the scientific publishing model. This study is not an endorsement for using Sci-Hub. The authors, their institutions, and publishers accept no responsibility on behalf of readers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.048 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.012 | 0.019 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.007 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".