Réflexions autour du projet de bibliographie des éditions lyonnaises du seizième siècle (BEL16)
Bibliographic record
Abstract
In November 2007, l’École nationale supérieure des sciences de l’information et des bibliothèques—France’s national school for information and librarianship—launched an ambitious project, following William Kemp’s proposal: establishing, in electronic form, an exhaustive, retrospective bibliography of books printed at Lyons during the sixteenth century. The implementation of this project was the object of numerous reflections, mostly upon the way the history of the book and the history of philology complement each other. Professional and disciplinary specificities concerned the identification of the types of users of such a base, the needs of these users, the norms regularly used, and the different levels of description considered to be necessary. This article recounts these conceptual progressions as they helped define bibliography in the twenty-first century. With precise comparisons to existing databases, and with concise and detailed definitions of methodology and issues, the author exposes the necessary decisions required of any bibliographic undertaking. Public, descriptions, corpus, standardization, and use are approached with reference to both conception and concept.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.026 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.021 | 0.023 |
| Science and technology studies | 0.004 | 0.006 |
| Scholarly communication | 0.018 | 0.006 |
| Open science | 0.001 | 0.006 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.011 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".