MétaCan
Menu
Back to cohort
Record W2930492841 · doi:10.5281/zenodo.1451753

OcOr : a Corpus of Occitan Oral Narratives

2018· dataset· en· W2930492841 on OpenAlexaff
Marianne Vergez-Couret, Janice Carruthers

Bibliographic record

VenueZenodo (CERN European Organization for Nuclear Research) · 2018
Typedataset
Languageen
FieldArts and Humanities
TopicMedieval European Literature and History
Canadian institutionsQueen's University
FundersEuropean Commission
KeywordsNarrativeLinguisticsArtHistoryLiteraturePhilosophy

Abstract

fetched live from OpenAlex

OcOr is a corpus of Occitan oral narratives. This corpus is one of the outputs of the project ExpressioNarration, financed by a Marie Sklodovska Curie Fellowship (2016-2018, n°655034). It includes three sub-corpora, constituted as follows: • OOT (Occitan, oral, traditional): stories drawn from fieldwork among native speakers in the Occitan domain, recorded by the COMDT (Conservatoire Occitan des Musiques et Danses Traditionnelles - http://www.comdt.org/), transcribed and digitised for the project by the researchers. • OWT (Occitan, written, traditional): published literary stories, digitised by and for the project by the researchers. These are stories collected from oral sources and produced in a publishable written version. • OOC (Occitan, oral, contemporary): stories recounted by contemporary artists, taken from existing recordings and two Toulouse storytelling events organised by the project in collaboration with the Institut d'Etudes Occitanes (IEO), in 2016. The stories were recorded during the events and subsequently transcribed and digitised by the researchers. The overall aim of the ExpressioNarration project was to use contemporary linguistic theory to explore the relationship between language and orality, with a specific focus on key temporal features of oral narrative in Occitan, including ‘tenses’, ‘connectives' and 'frame introducers'. These features were thus annotated in the three sub-corpora. All the sub-corpora are disseminated in XML format (TEI-P5) and PDF. Each story is available as an annotated XML document, an annotated PDF and a stripped PDF document. Full metadata appears in the Header of each XML document, with information on speakers (e.g. gender, age, place of origin, education, languages spoken), variety of Occitan (or dialect), authors/editorial information (in the case of OWT) and story-type when relevant (i.e. the Aarne Thompson category). For each sub-corpus, a user-friendly summary of this metadata is also available in an Excel spreadsheet: these are contained in the OcOr zipfile. The annotation system was designed by the researchers and is given in full in the Header of each XML document. For further information on the constitution of the corpus and discussion of the theoretical and methodological issues relating to data collection, digitisation and annotation, please read the following article in the journal <em>Corpus</em>, written by the researchers and entitled ‘Méthodologie pour la constitution d’un corpus comparatif de narration orale en Occitan : objectifs, défis, solutions’, available at: https://journals.openedition.org/corpus/3490.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Science and technology studies, Insufficient payload (model declined to judge)
Consensus categoriesInsufficient payload (model declined to judge)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.142
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0020.001
Scholarly communication0.0010.000
Open science0.0010.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.1100.010

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.049
GPT teacher head0.243
Teacher spread0.194 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2018
Admission routes1
Has abstractyes

Explore more

Same venueZenodo (CERN European Organization for Nuclear Research)Same topicMedieval European Literature and HistoryFrench-language works237,207