MétaCan
Menu
Retour à la cohorte
Enregistrement W4416414481 · doi:10.3897/biss.9.178687

Mobilising Legacy Georeferencing Efforts

2025· article· W4416414481 sur OpenAlexaboutno aff
Ashleigh Whittaker, Jack Plummer, Nicky Nicolson

Notice bibliographique

RevueBiodiversity Information Science and Standards · 2025
Typearticle
Langue
DomaineEnvironmental Science
ThématiqueSpecies Distribution and Climate Change
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésHerbariumGeoreferenceBiodiversityExtinction (optical mineralogy)WorkflowDistribution (mathematics)Representativeness heuristicContext (archaeology)

Résumé

récupéré en direct d'OpenAlex

As we progress towards a globally accessible natural history collection, the ways in which we digitally curate, share and use our data will inevitably change. Digital specimen records enable access by in-country experts, increase the opportunity for specimen enhancement by the scientific community, and improve the breadth of research to which the specimens contribute. Digitisation workflows capture an image often with minimal transcription; they do not include the further enhancement of specimen records, such as georeferencing - the addition of coordinates to text based locality information. Typically used to georeference herbarium specimens, the point radius method assigns a coordinate and a measurement of maximum uncertainty (Wieczorek and Chapman 2020). Some herbarium specimens that lack coordinates still cannot be georeferenced using this method due to poor label data or the unavailability of contextual information. Data derived from herbarium specimens form the basis of species distribution modelling, taxonomic research and IUCN extinction risk assessments. If a specimen record does not have coordinates, then it is often discarded early in the data mining process, highlighting the importance of data enhancement through georeferencing, which confirms the species occurrence in space and time. The Kunming-Montreal Global Biodiversity Framework targets help to guide the most important applications of collection data for biodiversity conservation; Target 4, for example, aims to halt species extinction and protect genetic diversity (Convention on Biological Diversity 2023). Extinction risk assessments of plant species are underpinned by distribution maps derived from preserved specimens and observation records with coordinate information either recorded at the time of collection or through subsequent georeferencing of the specimen. However, over half of all herbarium specimens on Global Biodiversity Information Facility (GBIF) do not have coordinates and so robust extinction risk assessments often require a georeferencing step before applying the IUCN’s criteria. Other data types contributing to extinction risk assessment, such as population size and trend are often lacking and, where available, subjective and restricted to an expert’s firsthand knowledge of the species, making the data untraceable for the wider community (Nic Lughadha et al. 2019, Willis et al. 2003). As highlighted in Bloom et al. (2018), the estimated distribution of a species differs depending on the origin of the coordinate data. For example, non-georeferenced records tend to inflate the estimated distribution of a species. Working with collaborators with regional geographical expertise, and using datasets following standardised protocols, will more likely result in accurate and precise georeferenced records. As herbarium specimen labels can contain qualitative context on collection localities, the process of georeferencing is subjective and so an indicator of confidence and the georeferencer's method is important for end user trust and usability of the record. As more herbarium specimens are digitised, there is a growing need to produce georeferenced locality data at scale.. Although there are robust tools that can be used to supplement manual georeferencing, georeferencing of all specimens in natural history collections currently lacking coordinates is not feasible to resource, as georeferencing is time and resource heavy. The Royal Botanic Gardens, Kew now contributes 5.8 million herbarium specimen records to GBIF, many of which will have between three and six duplicate specimens located in herbaria around the world. Enhancing these records with existing georeferencing efforts accumulated through the completion of several thousand extinction risk assessments will reduce duplicated effort and uncover georeferenced localities that would otherwise go undocumented. This dataset will also help in establishing protocols to apply when georeferencing plant collections in the future, particularly in data-poor tropical regions.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,021
score de la tête « metaresearch » (Gemma)0,066
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: aucune
GenreSignal candidat: Méthodes · Signal consensuel: Méthodes
Score de désaccord entre enseignants0,033
Score d'incertitude au seuil0,111

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0210,066
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0130,015
Études des sciences et des technologies0,0040,002
Communication savante0,0100,012
Science ouverte0,0050,016
Intégrité de la recherche0,0020,003
Charge utile insuffisante (le modèle a refusé de juger)0,0330,022

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,018
Tête enseignante GPT0,258
Écart entre enseignants0,240 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreMéthodes

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2025
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueBiodiversity Information Science and StandardsMême sujetSpecies Distribution and Climate ChangeTravaux en français237 207