MétaCan
Menu
Back to cohort
Record W1568140279

Proceedings of the Workshop on Multiword Expressions: Identifying and Exploiting Underlying Properties

2006· article· en· W1568140279 on OpenAlexaff
Begoña Villada Moirón, Diana McCarthy, Stefan Evert, Suzanne Stevenson

Bibliographic record

Venuenot available
Typearticle
Languageen
FieldComputer Science
TopicNatural Language Processing Techniques
Canadian institutionsUniversity of Toronto
Fundersnot available
KeywordsComputer scienceNatural language processingArtificial intelligencePresentation (obstetrics)LexiconFocus (optics)Automatic summarizationPrinciple of compositionalityQuestion answeringSentenceFrameNetRepresentation (politics)ParsingComputational linguisticsLinguistics
DOInot available

Abstract

fetched live from OpenAlex

This volume contains the papers accepted for presentation at the Workshop on Multiword Expressions: Identifying and Exploiting Underlying Properties. The workshop is endorsed by the Association for Computational Linguistics Special Interest Group on the Lexicon (SIGLEX) and is hosted in conjunction with the COLING/ACL 2006 on July 23rd, 2006 in Sydney, Australia. There has been a growing awareness in the NLP community of the problems that multiword expressions (MWEs) pose. Developments in areas such as machine translation, text summarization, paraphrasing, grammar development and parsing, information retrieval, and question answering (to mention a few) have acknowledged difficulties due to the idiosyncratic nature of multiword expressions. This workshop continues a tradition of ACL workshops on Collocations (2001) and Multiword Expressions (2003 and 2004). Its specific objective is to focus on the underlying properties of MWEs. The call for papers expressed our interest in several topics such as the definition of MWEs, properties of MWEs and their impact on NLP applications, representation and treatment of the different classes of MWEs, linguistic and psycholinguistic analyses of MWEs, evaluation of extraction techniques and the importance of (non-)compositionality. We received 23 submissions in total. Each submission was reviewed by (at least) three members of the program committee who not only judged each submission but also gave detailed comments to the authors. Among the received papers, 10 were selected for presentation at the workshop. After 3 papers have been withdrawn by their authors, seven papers are included in these proceedings. The intention of this workshop is to focus on some fundamental questions on the nature of MWEs. To do this we will allow plenty of time for discussion to pursue some of the interesting, open and difficult questions that MWEs raise. As well as a discussion period after each session of papers, we will be organising group discussions at the end of the workshop. These will focus on problems of defining, characterising and evaluating MWEs, given what we know about the range of phenomena that they encompass as well as any important questions that have arisen during the workshop.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.009
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Other · Consensus signal: none
Teacher disagreement score0.044
Threshold uncertainty score0.148

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0040.009
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0020.002
Science and technology studies0.0020.001
Scholarly communication0.0080.009
Open science0.0020.004
Research integrity0.0020.003
Insufficient payload (model declined to judge)0.0440.025

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.063
GPT teacher head0.290
Teacher spread0.227 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreOther

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations14
Published2006
Admission routes1
Has abstractyes

Explore more

Same topicNatural Language Processing TechniquesFrench-language works237,207