Encoding Collocations in DiCoInfo:From formal to user-friendly representations
Bibliographic record
Abstract
This contribution presents an online lexical database and shows how some of its data categories were converted to make them more accessible to users. The database is called DiCoInfo and contains English, French, and Spanish terms related to the fields of computing and the Internet. Entries are compiled according to the principles of Explanatory Combinatorial Lexicology, ECL (Mel’čuk et al. 1995). First, we present the basic structure of the entry, focussing on the encoding of collocations (based on lexical functions). Then, we show how the meaning of collocations can be described with natural language explanations, how actantial structures can be better reflected in these explanations, how users can browse collocations in order to find collocates that express specific meanings, and finally how they can search for translations of collocations. Our work tends to demonstrate that, even though the encoding performed by lexicographers is semi-formal and proves necessary for the new functionalities described, it can still lend itself to adaptations defined according to specific user needs, and ultimately to meet those needs. This seems to be confirmed by the preliminary results of a pilot study we conducted on the browsing of collocations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".