Replication Data for: Mapping the landscape of geospatial data citations
Notice bibliographique
Résumé
This data supports the paper entitled "Mapping the landscape of geospatial data citations". The dataset covers geospatial data-intensive research papers published between 2015-2018 retrieved using Scopus. The article's citations were assessed for data citation occurances, and coded using a data citation classification. Data were enhanced and linked to subject coverage and journal policy status information using Excel & SPSS. For more information about how the data were created and coded please review the 'Methodology' section of the paper. More information is provided below, including supplemental documentation and related publications. Abstract (paper) ABSTRACT Data citations, similar to article and other research citations, are important references to research data that underlie published research results. In support of open science directives, these citations must adhere to specific conventions in terms of consistency of both placement within an article, and the actual availability or access to research data. To better understand the level to which geospatial research data are currently cited, we undertook a study to analyse the rate of data citation within a set of data-intensive geospatial research articles. After analysing 1717 scholarly articles published between 2015 and 2018, we found that very few, or 78 (5%), meaningfully cited primary or secondary geospatial data sources in the cited references section of the article. Even fewer researchers, only 25 or 1.5%, were found to have cited data using a DOI. Given the relatively low data citation rate, a focus on contributing factors including barriers to citing geospatial data is needed. And while open sharing requirements for geospatial data may change over time, driving data citation as a result, understanding benchmarks for data citation for monitoring purposes is useful.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,001 |
| Science ouverte | 0,003 | 0,001 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».