MétaCan
Menu
Retour à la cohorte
Enregistrement W4303986904 · doi:10.3897/bdj.10.e86089

The disambiguation of people names in biological collections

2022· article· en· W4303986904 sur OpenAlexaff
Quentin Groom, Christian Bräuchler, Robert Cubey, Mathias Dillen, Pieter Huybrechts, Nicole Kearney, Niels Klazenga, Siobhan Leachman, Deborah Paul, Heather Rogers, Joaquim Santos, David Peter Shorthouse, Alison Vaughan, Sabine von Mering, Elspeth Haston

Notice bibliographique

RevueBiodiversity Data Journal · 2022
Typearticle
Langueen
DomaineBiochemistry, Genetics and Molecular Biology
ThématiqueGenetic diversity and population structure
Établissements canadiensAgriculture and Agri-Food CanadaMcGill University
Organismes subventionnairesHorizon 2020 Framework ProgrammeFonds Wetenschappelijk OnderzoekVlaamse regeringEuropean CommissionVlaamse OverheidEuropean Cooperation in Science and TechnologyNational Science Foundation
Mots-clésComputer scienceInformation retrievalNatural language processingWorld Wide Web

Résumé

récupéré en direct d'OpenAlex

Scientific collections have been built by people. For hundreds of years, people have collected, studied, identified, preserved, documented and curated collection specimens. Understanding who those people are is of interest to historians, but much more can be made of these data by other stakeholders once they have been linked to the people's identities and their biographies. Knowing who people are helps us attribute work correctly, validate data and understand the scientific contribution of people and institutions. We can evaluate the work they have done, the interests they have, the places they have worked and what they have created from the specimens they have collected. The problem is that all we know about most of the people associated with collections are their names written on specimens. Disambiguating these people is the challenge that this paper addresses. Disambiguation of people often proves difficult in isolation and can result in staff or researchers independently trying to determine the identity of specific individuals over and over again. By sharing biographical data and building an open, collectively maintained dataset with shared knowledge, expertise and resources, it is possible to collectively deduce the identities of individuals, aggregate biographical information for each person, reduce duplication of effort and share the information locally and globally. The authors of this paper aspire to disambiguate all person names efficiently and fully in all their variations across the entirety of the biological sciences, starting with collections. Towards that vision, this paper has three key aims: to improve the linking, validation, enhancement and valorisation of person-related information within and between collections, databases and publications; to suggest good practice for identifying people involved in biological collections; and to promote coordination amongst all stakeholders, including individuals, natural history collections, institutions, learned societies, government agencies and data aggregators.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,015
score de la tête « metaresearch » (Gemma)0,054
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: aucune
GenreSignal candidat: Méthodes · Signal consensuel: Méthodes
Score de désaccord entre enseignants0,018
Score d'incertitude au seuil0,082

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0150,054
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0180,016
Études des sciences et des technologies0,0060,005
Communication savante0,0070,014
Science ouverte0,0030,018
Intégrité de la recherche0,0020,003
Charge utile insuffisante (le modèle a refusé de juger)0,0020,003

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,044
Tête enseignante GPT0,253
Écart entre enseignants0,209 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreMéthodes

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations18
Publié2022
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueBiodiversity Data JournalMême sujetGenetic diversity and population structureTravaux en français237 207