Bibliographic record
Abstract
The linguistic fund, that is, the actual lexical inventory of a language, is always considerably larger than the sum total of the contents of dictionaries for that language. It corresponds to all the possibilities of derivation and compounding-with the associated word-formation rules. The exploitation of the latter. within machine-readable dictionaries should therefore allow a far more accurate coverage of the linguistic fund, by generating thousands of additional entries, some being more or less widely attested in their written form, some representing the set of virtual words generated by a productive rule that have not, for one reason or another, been recognized as existing words. Note that the boundary between the two subsets is not clearcut. The conditions which determine whether a generated form will belong to one or the other have not given rise to extensive studies, neither in linguistics nor in lexicology, in spite of their significance for a better understanding of lexical creativity and word-formation processes. The relevance of such phenomena seems to have been greatly underestimated-as is indicated by the fact that no studies have been fulfilled on the topic of words such as medico-legal and franco-quebecois. We will demonstrate and illustrate for the French language the shortcomings of the simple compiling method and how ignoring them has led to unnecessary complications in the field of electronic lexicography
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.003 | 0.001 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.353 | 0.140 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".