Bibliographic record
Abstract
Book banning has become a widespread method of global censorship. Governments are increasingly using this approach to control internet and technological resources. As is well known, censorship in libraries, and especially school and public libraries, is a subject of international and complex debate that touches several areas, from fundamental rights and freedom of expression to the social responsibility of public institutions and the professionals who work there, also the necessary representation of the existing diversity. Advocacy groups, like Amnesty International or library associations, have emerged to combat this threat to democracy, education, and progressive thinking. Countries such as China, Bangladesh, and Egypt commonly ban books to limit education and suppress vulnerable populations1. In 2022, the American Library Association received a record-breaking 1,269 requests to restrict library access. This surge, nearly doubling the previous year’s figures, underscores concern about intellectual freedom and diverse literary content. Within the 2,571 titles targeted for book censorship cases, some titles faced intense scrutiny2. Nunia Ferran Ferrar Miquel Centelles Lluis Agusti In April 2023, volunteers from Botswana, Brazil, Canada, Mexico, Catalonia, and the United States launched the #EveryBookItsReader initiative3. Their goal was to improve content related to books, literary works, and oral traditions on various Wikimedia platforms, including Wikipedia, Wikidata, Wikicommons, Wikiquotes, Wikibooks, and Wikisource. This collaborative initiative, from various countries, recurs annually throughout April each year, aligning with World Book Day on April 23, which originated in Catalonia, Spain, as the “Day of Books and Roses.” Anyone, especially librarians, can participate.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.021 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.005 | 0.009 |
| Scholarly communication | 0.012 | 0.012 |
| Open science | 0.002 | 0.006 |
| Research integrity | 0.003 | 0.005 |
| Insufficient payload (model declined to judge) | 0.024 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".