MétaCan
Menu
Back to cohort
Record W4322622999 · doi:10.5325/libraries.7.1.0105

Along Came Google: A History of Library Digitization

2023· article· en· W4322622999 on OpenAlexaboutno aff
Eric Novotny

Bibliographic record

VenueLibraries Culture History and Society · 2023
Typearticle
Languageen
FieldComputer Science
TopicLibrary Collection Development and Digital Resources
Canadian institutionsnot available
Fundersnot available
KeywordsDigitizationPopularityAppealThe InternetLibrary scienceWorld Wide WebDigital libraryPlan (archaeology)Political scienceMedia studiesHistorySociologyComputer scienceLawArt

Abstract

fetched live from OpenAlex

In the early 2000s I confidently declared that even as e-books gained popularity, libraries would forever retain their role as repositories of the knowledge they collected and stewarded over many years. After all, who would invest billions to digitize old monographs with tremendous scholarly value but limited commercial appeal? Then, as the authors say, “Along Came Google.” In an engaging and fast-paced account, the authors describe how Google abruptly transformed the scholarly landscape, digitizing books at an unprecedented scale. The authors are well positioned to tell this tale. As members and leaders of organizations including Ithaka S+R, the Library of Congress, and the Council on Library and Information Resources, they took an active role in conversations around the evolving information ecosystem. Their personal knowledge is enriched with the addition of nineteen interviews with key participants from the Internet Archive, Harvard, Michigan, HathiTrust, and others.Before the dramatic Google announcement, the story begins with a brief history of library collaborations, including the famous Farmington Plan adopted after World War II and the creation of OCLC in the 1970s. While not every effort was successful, the pre-digital efforts established a precedent for national library networks and cooperation. In chapter 2, “The Dreamers,” the authors outline the many nascent efforts that sprang up in the 1990s and early 2000s to harness the potential of the internet, from the Million Books Project to Making of America and JSTOR. The authors conclude that, while innovative and laudable, “none of these created a universal library or transformed the nature of the research library” (72). Additionally, the pace was excruciatingly slow. When approached by Google, the University of Michigan estimated it would take more than a thousand years to digitize the library’s eleven million volumes using existing approaches and with the current budget. Google promised to scan the entire collection in six years. (78).The subsequent chapters are the heart of the book, detailing the stunning entrance of Google into the field, the diverse reactions, and potential alternatives. A wide range of issues and concerns are discussed, reflecting the perspectives of various stakeholders, including librarians, publishers, and technologists. The participant interviews provide a personal perspective, including some less-than-flattering accounts of perceived jealousy and positions seemingly motivated by the loss of professional or institutional prestige. Throughout, the authors largely avoid taking sides, offering the arguments advanced by each party in an evenhanded manner. They capture the emotion of the moment, including fears of cultural imperialism, homogenization of the nascent digital library, and entrusting the cultural record to a single commercial entity. These concerns spawned efforts such as the Open Content Alliance, a short-lived collaboration between Yahoo, the Internet Archive, the University of California, the University of Toronto, and others. Publisher and author concerns over copyright, privacy, and piracy fueled the lawsuits that effectively killed the dream of the universal digital library.A pleasure of this work are the many fascinating details—for example, the anecdote that early scanning efforts were stymied by distortion from layers of dust on print books. Even those of us who lived through the era will learn new things. I never knew, or no longer remember, that despite their later objections publishers initially embraced the original Google Print project. They were motivated to ally with Google in part to gain leverage against the growing power of Amazon in book discovery and sales.While riveting throughout, there are some missed opportunities. The account offered is largely descriptive—the authors do not try to advance a thesis or engage much with the existing scholarly literature. Library historians will lament the lack of a bibliography, and there is little information about the interviews other than the date they were conducted. It is not clear if the interviews have been archived or are otherwise available. Clarifying this on the website or in a future edition would be a valuable addition for those who wish to further mine the interviews for insights.Despite these critiques, this concise, engaging work will be of interest to library historians and a general library-informed audience (nonlibrarians may struggle to keep track of the many organizations and their acronyms: OCLC, CRL, ACRL, LC). A few recent monographs cover some of the same ground from a different perspective. Google Rules: The History and Future of Copyright under the Influence of Google (Oxford University Press, 2020) focuses on Google’s interpretation of copyright law and its implications for the public good. The Politics of Mass Digitization (MIT Press, 2018) examines some of the same conflicting public and private interests involved in the preservation of cultural memory at a large scale. While there is some overlap in the issues discussed, Along Came Google is unique in placing librarians and library-allied organizations at the center of the conversation. Library historians will surely appreciate hearing directly from key participants at a pivotal moment in library history.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.011
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesScholarly communication
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.975
Threshold uncertainty score0.138

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0040.011
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0100.023
Science and technology studies0.0160.021
Scholarly communication0.0250.025
Open science0.0020.014
Research integrity0.0040.006
Insufficient payload (model declined to judge)0.0220.005

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.014
GPT teacher head0.167
Teacher spread0.153 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueLibraries Culture History and SocietySame topicLibrary Collection Development and Digital ResourcesFrench-language works237,207