MétaCan
Menu
Retour à la cohorte
Enregistrement W2995542851 · doi:10.18438/eblip29606

Weak Correlation Between Circulation and Citation Numbers Suggests that both Data Points should be Considered when Deselecting Print Monographs

2019· article· en· W2995542851 sur OpenAlexvenueno aff
Melissa J. Goertzen

Notice bibliographique

RevueEvidence Based Library and Information Practice · 2019
Typearticle
Langueen
DomaineArts and Humanities
ThématiqueAcademic Writing and Publishing
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésCitationLibrary scienceComputer scienceCitation analysisRank correlationCirculation (fluid dynamics)Information retrieval

Résumé

récupéré en direct d'OpenAlex

A Review of: White, B. (2017). Citations and circulation counts: Data sources for monograph deselection in research library collections. College & Research Libraries, 78(1), 53 – 65. https://doi.org/10.5860/crl.78.1.53 Abstract Objective – To facilitate evidence-based deselection of print monographs, this study examines to what extent there are correlations between circulation data (past and future usage) and between the borrowing and citation of print monographs. Design – Collections assessment project that used a variety of data sources and techniques, including Spearman’s rank correlation coefficient, statistical analysis, and the analysis of circulation data, last-use dates, and citation data. Setting – An academic library in New Zealand. Subjects – Two ranges of books were chosen for the study: 591 (Specific Topics in Zoology) and 324 (The Political Process). From these ranges, monographs published prior to 2001 were selected as the study sample. Methods – This project relied on two data sources: circulation data from the Library’s ILS and citation data from Scopus. All data was downloaded to an Excel spreadsheet in preparation for analysis. The researcher examined call numbers, authors and editors, titles and subtitles, publication dates, circulation counts, dates of last check-in, total number of citations, number of citations from publications released in 2010 and on, and number of citations from institution-affiliated documents. Renewal data was omitted, as it did not provide evidence of additional instances of use. Where multiple copies of a specific title appeared in the data set, the researcher totalled all circulations and recorded the most recent check-in date. The researcher found that some titles in the study sample were generic and it was impossible to determine if citation data from Scopus linked to the monograph in the library collection. These titles were eliminated from the study. Once data collection was complete, the researcher calculated two additional data elements: the number of months since the last check-in date and the number of citations from items published before 2010. Data in the Excel spreadsheet was analyzed using Spearman’s rank correlation coefficient to determine the relationship between past and future usage and between circulation and citation data. Main Results – Findings indicated that circulation and citation data are highly skewed. Many monographs in the study sample had never been borrowed and had few citations, while a small number of “celebrity titles” were borrowed or cited at a much higher rate than other monographs in the same classification. Further, results indicated that historic circulation numbers are imperfect predictors of future probability that a book will be borrowed. When taking a high-level view of the collection, highly circulated books tend to be borrowed more often than average. However, when examining monographs at the title level, high circulation is more of a probability instead of a robust indicator. An investigation of whether historic citation counts serve as an indicator of future citation followed previously established trends: monographs not heavily cited in the past are less likely to be cited in the future. Findings also found a weak correlation between local-institution monograph citation counts and total citation counts. Finally, the results demonstrated a weak correlation between circulation and citation data. As a group, well-cited books are borrowed more often than others, but at the individual title level, the effect is too random for either data set to predict the other in a reliable way. As such, circulation data and citation data can not be used as a proxy for each other. Conclusion – Neither circulation nor citation data can stand as full proxies of the value of a title. However, both provide information that reflects the status of a title within the scholarly community. In this environment, citation data should be considered equally with circulation figures. Both data points measure different phenomena and the weak correlation between them suggests that both are required to inform decisions about deselecting print monographs.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,060
score de la tête « metaresearch » (Gemma)0,400
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche, Bibliométrie
Catégories consensuellesaucune
DomaineSignal candidat: Évaluation · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: Observationnel
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,965
Score d'incertitude au seuil0,319

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0600,400
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0350,055
Études des sciences et des technologies0,0020,004
Communication savante0,0100,018
Science ouverte0,0030,004
Intégrité de la recherche0,0010,002
Charge utile insuffisante (le modèle a refusé de juger)0,0190,005

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,097
Tête enseignante GPT0,286
Écart entre enseignants0,190 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeObservationnel
DomaineÉvaluation
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2019
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueEvidence Based Library and Information PracticeMême sujetAcademic Writing and PublishingTravaux en français237 207