Metadata as Data: Exploring Ethical Metadata Sharing and Access for Indigenous Resources Through OCAP Principles
Bibliographic record
Abstract
Metadata is often defined as “data about data”, and although practitioners and scholars often broaden that definition, there may be value in approaching metadata as a type of data when addressing questions of ethical sharing and access. In this conceptual paper I review the challenges of ethical metadata practice for Indigenous resources, and explore the potential of the OCAP: Ownership, Control, Access, and Possession framework to act as a common language that Indigenous communities and metadata scholars and practitioners can use to engage in meaningful conversations about ethical metadata access and sharing.Les métadonnées sont souvent définies comme des «données sur les données» et bien que les professionnels et les chercheurs élargissent souvent cette définition, il peut être valable d'aborder les métadonnées comme un type de données lorsqu'on aborde les questions de partage et d'accès éthique. Dans cet article conceptuel, je passe en revue les défis de la pratique éthique des métadonnées pour les ressources autochtones et explore le potentiel du cadre PCAP: propriété, contrôle, accès et possession, pour servir de langage commun aux communautés autochtones et aux spécialistes et aux professionnels des métadonnées afin d’engager des conversations significatives sur l'accès et le partage des métadonnées éthiques.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.038 | 0.052 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.007 |
| Science and technology studies | 0.010 | 0.053 |
| Scholarly communication | 0.022 | 0.032 |
| Open science | 0.004 | 0.019 |
| Research integrity | 0.004 | 0.005 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".