MétaCan
Menu
Back to cohort
Record W133184457

Basic Concepts of Lexical Resource Semantics

2008· article· en· W133184457 on OpenAlexaboutno aff
Frank Richter, Manfred Sailer

Bibliographic record

Venuenot available
Typearticle
Languageen
FieldComputer Science
TopicNatural Language Processing Techniques
Canadian institutionsnot available
Fundersnot available
KeywordsComputer scienceUnderspecificationHead-driven phrase structure grammarNatural language processingGenerative grammarArtificial intelligenceLinguisticsSemantics (computer science)Programming language
DOInot available

Abstract

fetched live from OpenAlex

Semanticists use a range of highly expressive logical languages to characterize the meaning of natural language expressions. The logical languages are usually taken from an inventory of standard mathematical systems, with which generative linguists are familiar. They are, thus, easily accessible beyond the borders of a given framework such as Categorial Grammar, Lexical Functional Grammar, or Government and Binding Theory. Linguists working in the HPSG framework, on the other hand, often use rather idiosyncratic and specialized semantic representations. Their choice is sometimes motivated by computational applications in parsing, generation, or machine translation. Naturally, the intended areas of application influence the design of semantic representations. A typical property of semantic representations in HPSG that is concerned with computational applications is underspecification, and other properties come from the particular unification or constraint solving algorithms that are used for processing grammars. While the resulting semantic representations have properties that are motivated by, and are adequate for, certain practical applications, their relationship to standard languages is sometimes left on an intuitive level. In addition, the theoretical and ontological status of the semantic representations is often neglected. This vagueness tends to be unsatisfying to many semanticists, and the idiosyncratic shape of the semantic representations confines their usage to HPSG. Since their entire architecture is highly dependent on HPSG, hardly anyone working outside of that framework is interested in studying them. With our work on Lexical Resource Semantics (LRS), we want to contribute to the investigation of a number of important theoretical issues surrounding semantic representations and possible ways of underspecification. While LRS is formulated in a constraint-based grammar environment and takes advantage of the tight connection between syntax proper and logical representations that can easily be achieved in HPSG, the architecture of LRS remains independent from that framework, and combines attractive properties of various semantic systems. We will explore the types of semantic frameworks which can be specified in Relational Speciate Re-entrant Language (RSRL), the formalism that we choose to express our grammar principles, and we evaluate the semantic frameworks with respect to their potential for providing empirically satisfactory analyses of typical problems in the semantics of natural languages. In LRS, we want to synthesize a flexible meta-theory that can be applied to different interesting semantic representation languages and make computing with them feasible. We will start our investigation with a standard semantic representation language from natural language semantics, Ty2 (Gallin, 1975). We are well aware of the debate about the appropriateness of Montagovian-style intensionality for the analysis of natural language semantics, but we believe that it is best to start with a semantic representation that most generative linguists are familiar with. As will become clear in the course of our discussion, the LRS framework is a meta-theory of semantic representations, and we believe that it is suitable for various representation languages, This paper can be regarded as a snapshot of our work on LRS. It was written as material for the authors’ course Constraint-based Combinatorial Semantics at the 15th European Summer School in Logic, Language and Information in Vienna in August 2003. It is meant as background reading and as a basis of discussion for our class. Its air of a work in progress is deliberate. As we see continued development in LRS, its application to a wider range of languages and empirical phenomena, and especially the implementation of an LRS module as a component of the TRALE grammar development environment; we expect further modifications and refinements to the theory. The implementation of LRS is realized in collaboration with Gerald Penn of the University of Toronto. We would like to thank Carmella Payne for proofreading various versions of this paper.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.005
metaresearch head score (Gemma)0.011
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.015
Threshold uncertainty score0.051

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0050.011
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0070.007
Science and technology studies0.0030.016
Scholarly communication0.0110.027
Open science0.0050.006
Research integrity0.0040.006
Insufficient payload (model declined to judge)0.0150.006

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.021
GPT teacher head0.287
Teacher spread0.267 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designTheoretical or conceptual
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations40
Published2008
Admission routes1
Has abstractyes

Explore more

Same topicNatural Language Processing TechniquesFrench-language works237,207