MétaCan
Menu
Retour à la cohorte
Enregistrement W4405540286 · doi:10.18438/eblip30614

Academic Libraries can Expand Institutional Repository Holdings with Gold Open Access Publications Collected Through Web Scraping

2024· article· en· W4405540286 sur OpenAlexvenueno aff
Kristy Hancock

Notice bibliographique

RevueEvidence Based Library and Information Practice · 2024
Typearticle
Langueen
DomaineComputer Science
ThématiqueResearch Data Management Practices
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésWorld Wide WebComputer scienceLibrary scienceBusinessInternet privacy

Résumé

récupéré en direct d'OpenAlex

A Review of: Clark, B. (2023). Proactive institutional repository collection development techniques: Archiving gold open access articles and metadata retrieved with web scraping. Journal of Library Administration, 63(6), 743–765. https://doi.org/10.1080/01930826.2023.2240190 Objective – To describe a method for collecting gold open access publications from the web and packaging them for batch deposit in an institutional repository. The goal of this project is to expand institutional repository holdings and increase the comprehensiveness of the collection with gold open access content. Design – Web scraping and analysis of institutional repository usage metrics. Setting – A library at a public doctoral university with very high research activity in Alabama, United States. Subjects – Articles and metadata from the Multidisciplinary Digital Publishing Institute (MDPI) website and the Sponsoring Consortium for Open Access Publishing in Participle Physics (SCOAP3) repository. MDPI is an open access publisher of over 400 journals spanning all disciplines. All articles published in MDPI journals are made freely and immediately accessible on the MDPI website. SCOAP3 is a global partnership of libraries, funding agencies, and research centers that support open access publishing in the field of high-energy physics. The SCOAP3 repository contains research funded by the organization and published in open access journals. Methods – The MDPI website and SCOAP3 repository were selected because they contained a substantial amount of scholarship by University of Alabama affiliates. On the MDPI website, an author affiliation search across all journals retrieved University of Alabama publications. The Python library Beautiful Soup was used with the parser package lxml to collect articles and metadata. The first script iterated through the pages of search results, downloaded article PDFs, and wrote abstract page URLs to a text file. The second script collected metadata by iterating through the text file of abstract page URLs, parsing the HTML of each URL, and writing Dublin Core metadata to a CSV file. Articles already archived in the institutional repository were removed from the CSV file, and the remaining metadata were reviewed for errors. To pair each PDF with the correct metadata, the file names of all PDFs were added to the CSV file. Article PDFs and the metadata file were packaged using the DSpace CSV Archive and batch deposited in the University of Alabama’s institutional repository. In SCOAP3, an author affiliation search retrieved University of Alabama publications. The browser automation software Selenium was used to collect articles and metadata. The first script iterated through the pages of search results and wrote article record page URLs to a text file. The second script downloaded article PDFs and extracted DOIs to use for PDF file names. The third script collected metadata by using the article record page URLs to query the SCOAP3 metadata harvesting API and writing MARCXML metadata to a CSV file. To pair each PDF with the correct metadata, the DOI column in the CSV file was duplicated, and the “.pdf” extension added to each DOI. The metadata in the CSV file was reviewed for errors, and citations and keywords were added manually. Articles and the metadata file were packaged and deposited using the MDPI method. The impact of SCOAP3 content on institutional repository downloads from the physics and astronomy collection was measured in the 100 days preceding and following the deposits. Main Results – 1,005 articles with corresponding metadata were collected from the MDPI website and SCOAP3 repository. After removing duplicate articles that were already archived in the University of Alabama institutional repository, 937 articles (272 from MDPI, 665 from SCOAP3) were deposited. The amount of faculty research available in the institutional repository increased from 1,639 articles before the project to 2,513 articles, or 37.3%. 678 articles were added to the physics and astronomy collection, which reflects the fact that most of the deposited articles were from a subject repository. The rest of the deposited articles were from MDPI and spanned various disciplines. The next best represented collections were civil, construction, and environmental engineering (26 articles); biological sciences (26 articles); electrical and computer engineering (24 articles); and geography (22 articles). The SCOAP3 articles also contributed to a significant increase in downloads from the physics and astronomy collection. Total downloads increased from 5,765 in the 100 days preceding the deposits to 7,243 in the 100 days following the deposits, with SCOAP3 articles representing 3,421 downloads, or 47.2%. Conclusion – This project was successful in proactively increasing the amount of scholarship in the institutional repository without faculty or researcher participation. This semi-automated workflow requires considerable technical skills but is manageable for one person. Since the articles and metadata were freely accessible and issued under permissive Creative Commons licenses, there was no need to consult publisher self-archiving policies or solicit permission to copy the articles to the institutional repository. This project did not make any research openly accessible that was otherwise unavailable or behind a paywall, but the added publications contribute to making the institution’s scholarly record more complete. This approach may be particularly helpful for academic library staff looking to build the holdings of a brand-new institutional repository, or for those dealing with an underpopulated institutional repository due to low self-archiving rates. Additional repositories containing a substantial amount of University of Alabama scholarship will be identified and considered for web scraping, to continue expanding the institutional repository holdings. The MDPI website and SCOAP3 repository will also be re-scraped in the future for research added since this project.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,069
score de la tête « metaresearch » (Gemma)0,159
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesCommunication savante, Science ouverte
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: aucune
GenreSignal candidat: Empirique · Signal consensuel: aucune
Score de désaccord entre enseignants0,994
Score d'incertitude au seuil0,365

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0690,159
Méta-épidémiologie (sens strict)0,0010,002
Méta-épidémiologie (sens large)0,0020,003
Bibliométrie0,0260,033
Études des sciences et des technologies0,0080,004
Communication savante0,0250,029
Science ouverte0,0060,032
Intégrité de la recherche0,0030,003
Charge utile insuffisante (le modèle a refusé de juger)0,0450,050

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,098
Tête enseignante GPT0,370
Écart entre enseignants0,272 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeSans objet
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations1
Publié2024
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueEvidence Based Library and Information PracticeMême sujetResearch Data Management PracticesTravaux en français237 207