Artificial Intelligence and the Book Industry. White Paper
Bibliographic record
Abstract
Artificial intelligence (AI) in the book world is a reality. Indeed, it is not reserved for sales platforms or medical applications. AI can assist writing, accompany editorial work or help the bookseller. It can respond to crying needs; despite its obvious limitations, it can be used to consider new applications in the book chain, which are the subject of specific recommendations here. This white paper, written by two specialists in the field of books and artificial intelligence, aims to identify ways to put AI at the service of the many links in the book world.<br> “Planning for this cultural niche’s immediate future must be done and specific actions must be undertaken in order to establish new methods and models. This white paper will outline a possible course of action: the idea of a concerted effort by book industry actors in the use of AI.”<br> This consultation is called for by a number of experts, who testify in this White Paper of the stakes specific to the current cultural context threatened by the giants of commerce : “Although use of AI calls for constant vigilance, it seems important that actors in the book industry pay close attention to these technological advances, as much to the potential disruptions as to the possible benefits they could entail.” (Virginie Clayssen, Éditis / Digital committee of the French Publishers Association)<br> Thus, “the key to introducing AI, thought as augmented intelligence, to different links in the book chain is undoubtedly exploitation of data that is already available and that the competition does not possess”.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.007 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".