MétaCan
Menu
Back to cohort
Record W7052603321

A Self-Paced Reading study on Processing Constructions with different degrees of Compositionality

2022· article· en· W7052603321 on OpenAlexaboutno aff

Bibliographic record

VenueCINECA IRIS Institutial research information system (University of Pisa) · 2022
Typearticle
Languageen
FieldEngineering
TopicPlasma Diagnostics and Applications
Canadian institutionsnot available
Fundersnot available
KeywordsPrinciple of compositionalityBigramSentencePhraseReading (process)Sentence processingSet (abstract data type)Phrase structure rules
DOInot available

Abstract

fetched live from OpenAlex

Introduction. Our research aims at challenging the classical principle of compositionality in \nsentence processing. While the principle of compositionality is traditionally considered as the \nprimary mean of explaining linguistic productivity, we argue it is just a default option within a \nmore complex scenario, where a series of noncompositional mechanisms (analogy with stored \nexemplars, shallow processing, activation of a network of mutual expectations, etc.) can be \nused in processing not only formulaic language expressions but a larger set of expressions \nwith a completely transparent meaning. Psycholinguistic literature has studied two \nnoncompositional cases: idiomatic processing, associated with faster reading time [1] and a \nmore positive electric signal in brain activity [2] to transparent phrases, and frequency effects, \ni.e., multiword sequences are usually read faster than comparable sequences of lesser \nfrequency [3,4]. We assume that facilitation effects are not limited to formulaic expressions \nbut also occur when processing highly prototypical and yet compositional phrases. To test this \nhypothesis, we implemented a Self-Paced Reading experiment to compare the Reading \nTimes (RTs) of argument constructions with a different degree in compositionality: idiomatic \nexpressions (ID), compositional highly frequent expressions (HF), and compositional low- \nfrequent expressions (LF). To the best of our knowledge, no previous work had compared \nboth idioms and frequent constructions, except for [5]. Given the previous literature, we \nhypothesized that RTs are longer for compositional sentences than for idiomatic sentences, \nand RTs are longer for infrequent phrases than for frequent ones. \n \nDesign. We selected 48 idiomatic VERB+determinant+NOUN phrases and corresponding \nhigh-frequency and low-frequency bigrams with the same verb. Each stimulus consisted of a \ncontext sentence presented for the participant to read in one instance and a sentence with the \ntarget phrase embedded (Table1), displayed word-by-word using the moving-window SPR \nparadigm [6]. The stimuli were split into three counterbalanced lists randomly initialized at \neach time. The experiment was delivered remotely, and participants were recruited using \nProlific. We collected responses for 90 subjects from the United States and Canada, all self - \nreported L1 speakers of English aged between 18 and 50. \n \nData Analysis. We removed the outliers and examined the RTs of the last word of phrases \nusing linear mixed models. Condition, Age, WordLength, VerbFrequency, and PositionInList \nwere entered in the models as fixed effects; Subject and Item were treated as random effects \nwith a by-subject random slope for BigramFrequency. RTs' difference between ID and HF \nturned out to be not statistically significant (Table2), while it was statistically different between \nID and LF. Changing the reference level with HF condition, there is still a smaller statistical \ndifference between HF and LF. Moreover, we observed that 1) older adults are slower than \nyounger speakers, and 2) RTs at the end of the experiment are faster than at the beginning. \n \nDiscussion. Analysis reveals no difference between processing the figurative meaning of \nidioms and the compositional one of HF; there are facilitation effects in comprehension of both \nexpressions. Even if this observed measure cannot say what is happening at the brain level, \nit opens to a broad discussion about underlying mechanisms in language processing. It may \nsupport the hypothesis that HF expressions are stored as unanalyzed wholes and directly \nretrieved once recognized as idioms, following usage-based models [7,8]. An alternative \nexplanation is the existence of a co-activated network of representations operating together \nwith analogy-based mechanisms leading to sentence meaning construction and working side \nby side with classical compositional ones. RTs for infrequent phrases were significantly \nslower, even if the advantage was relatively small. We presume that information introduced in \ncontext sentences reduces the effort to interpret unpredictable expressions. \n \n[1] Conklin, K., & Schmitt, N. (2008). Formulaic sequences: Are they processed more quickly than \nnonf ormulaic language by native and nonnative speakers?. Applied linguistics, 29(1), 72-89. \n[2] Vespignani, F., Canal, P., Molinaro, N., Fonda, S., & Cacciari, C. (2010). Predictive mechanisms in \nidiom comprehension. Journal of Cognitive Neuroscience, 22(8), 1682-1700. \n[3] Arnon, I., & Snider, N. (2010). More than words: Frequency ef f ects f or multi -word phrases. Journal \nof memory and language, 62(1), 67-82. \n[4] Tremblay, A., Derwing, B., Libben, G., & Westbury, C. (2011). Processing advantages of lexical \nbundles: Evidence f rom self ‐paced reading and sentence recall tasks. Language learning, 61(2), 569- \n613. \n[5] Jolsvai, H., McCauley, S. M., & Christiansen, M. H. (2020). Meaningf ulness beats f requency in \nmultiword chunk processing. Cognitive Science, 44(10). \n[6] Just, M. A., Carpenter, P. A., & Woolley, J. D. (1982). Paradigms and processes in reading \ncomprehension. Journal of experimental psychology: General, 111(2), 228–238. \n[7] Goldberg, A. E. (2006). Constructions at work: The nature of generalization in language. Oxf ord \nUniversity Press on Demand. \n[8] Bybee, J. (2010). Language, usage and cognition. Cambridge University Press. \n[9] Bannard, C. and D. Matthews (2008). Stored word sequences in language learning: The ef f ect of \nf amiliarity on children’s repetition of f our-word combinations. Psychological science 19.3, 241–248.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.004
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.007
Threshold uncertainty score0.023

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.004
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.001
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0070.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.038
GPT teacher head0.261
Teacher spread0.223 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueCINECA IRIS Institutial research information system (University of Pisa)Same topicPlasma Diagnostics and ApplicationsFrench-language works237,207