A Classification of Digital Emergence: A Critical Approach to the Production of Digital Objects in Special Collections
Bibliographic record
Abstract
This paper examines the infrastructure of digital libraries and teases out the subtle ways their formation and construction is a digital extension and representation of the social, political, and institutional circumstances by which they are created. Building off lessons learned from UCLA Library Special Collections as a case study site, this paper proposes a classification of digital emergence that provides more transparency about how digital surrogates come to exist in digital libraries and how we can use this information to better contextualize the importance of these surrogates within academic library services. The discussion then situates digital libraries as medial interfacing infrastructures that are fundamentally non-neutral social apparatuses that disappear in the course of daily use. Marcuse’s notion of technological rationality is incorporated to illustrate the extent to which technological infrastructures influence and reformulate the way we understand the research process using special collections and archives, and how these infrastructures can function as a mechanism for information control. Finally, Bowker and Star’s text, Sorting Things Out: Classification and Its Consequences, briefly illustrates how librarians can contextualize the emergence of digital objects, and how this context, and the concomitant technological biases, can be methodologically brought to light using infrastructural inversion.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.003 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".