MétaCan
Menu
Back to cohort
Record W7004479988

"0" Stories

2014· other· en· W7004479988 on OpenAlexaboutno aff

Bibliographic record

VenueSeoul National University Open Repository (Seoul National University) · 2014
Typeother
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicCell Image Analysis Techniques
Canadian institutionsnot available
Fundersnot available
KeywordsRelevance (law)Set (abstract data type)Field (mathematics)Simple (philosophy)LexicographyLexical itemClass (philosophy)
DOInot available

Abstract

fetched live from OpenAlex

The linguistic fund, that is, the actual lexical inventory of a language, is always considerably larger than the sum total of the contents of dictionaries for that language. It corresponds to all the possibilities of derivation and compounding-with the associated word-formation rules. The exploitation of the latter. within machine-readable dictionaries should therefore allow a far more accurate coverage of the linguistic fund, by generating thousands of additional entries, some being more or less widely attested in their written form, some representing the set of virtual words generated by a productive rule that have not, for one reason or another, been recognized as existing words. Note that the boundary between the two subsets is not clearcut. The conditions which determine whether a generated form will belong to one or the other have not given rise to extensive studies, neither in linguistics nor in lexicology, in spite of their significance for a better understanding of lexical creativity and word-formation processes. The relevance of such phenomena seems to have been greatly underestimated-as is indicated by the fact that no studies have been fulfilled on the topic of words such as medico-legal and franco-quebecois. We will demonstrate and illustrate for the French language the shortcomings of the simple compiling method and how ignoring them has led to unnecessary complications in the field of electronic lexicography

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.003
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesInsufficient payload (model declined to judge)
Consensus categoriesInsufficient payload (model declined to judge)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Other · Consensus signal: Other
Teacher disagreement score0.647
Threshold uncertainty score0.923

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.003
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0010.001
Science and technology studies0.0030.001
Scholarly communication0.0040.003
Open science0.0010.003
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.3530.140

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.011
GPT teacher head0.246
Teacher spread0.235 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designNot applicable
Domainnot available
GenreOther

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2014
Admission routes1
Has abstractyes

Explore more

Same venueSeoul National University Open Repository (Seoul National University)Same topicCell Image Analysis TechniquesFrench-language works237,207