MétaCan
Menu
Back to cohort
Record W4288060620 · doi:10.18357/kula.221

Re-purposing Excavation Database Content as Paradata

2022· article· en· W4288060620 on OpenAlexvenueno aff
Lisa Börjesson, Olle Sköld, Zanna Friberg, Daniel Löwenborg, Gı́sli Pálsson, Isto Huvila

Bibliographic record

VenueKULA knowledge creation dissemination and preservation studies · 2022
Typearticle
Languageen
FieldComputer Science
TopicNatural Language Processing Techniques
Canadian institutionsnot available
FundersRiksbankens JubileumsfondEuropean Commission
KeywordsMetadataComputer scienceDocumentationField (mathematics)Identification (biology)StructuringProcess (computing)Data scienceExcavationReading (process)Plan (archaeology)ArchaeologyWorld Wide Web

Abstract

fetched live from OpenAlex

Although data reusers request information about how research data was created and curated, this information is often non-existent or only briefly covered in data descriptions. The need for such contextual information is particularly critical in fields like archaeology, where old legacy data created during different time periods and through varying methodological framings and fieldwork documentation practices retains its value as an important information source. This article explores the presence of contextual information in archaeological data with a specific focus on data provenance and processing information, i.e., paradata. The purpose of the article is to identify and explicate types of paradata in field observation documentation. The method used is an explorative close reading of field data from an archaeological excavation enriched with geographical metadata. The analysis covers technical and epistemological challenges and opportunities in paradata identification, and discusses the possibility of using identified paradata in data descriptions and for data reliability assessments. Results show that it is possible to identify both knowledge organisation paradata (KOP) relating to data structuring and knowledge-making paradata (KMP) relating to fieldwork methods and interpretative processes. However, while the data contains many traces of the research process, there is an uneven and, in some categories, low level of structure and systematicity that complicates automated metadata and paradata identification and extraction. The results show a need to broaden the understanding of how structure and systematicity are used and how they impact research data in archaeology and in comparable field sciences. The insights into how a dataset’s KOP and KMP can be read is also a methodological contribution to data literacy research and practice development. On a repository level, the results underline the need to include paradata about dataset creation, purpose, terminology, dataset internal and external relations, and eventual data colloquialisms that require explanation to reusers.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.024
metaresearch head score (Gemma)0.088
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesScholarly communication
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.987
Threshold uncertainty score0.127

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0240.088
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0170.018
Science and technology studies0.0030.003
Scholarly communication0.0130.013
Open science0.0030.008
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0040.004

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.098
GPT teacher head0.392
Teacher spread0.294 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations35
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueKULA knowledge creation dissemination and preservation studiesSame topicNatural Language Processing TechniquesFrench-language works237,207