A Self-Paced Reading study on Processing Constructions with different degrees of Compositionality
Notice bibliographique
Résumé
Introduction. Our research aims at challenging the classical principle of compositionality in \nsentence processing. While the principle of compositionality is traditionally considered as the \nprimary mean of explaining linguistic productivity, we argue it is just a default option within a \nmore complex scenario, where a series of noncompositional mechanisms (analogy with stored \nexemplars, shallow processing, activation of a network of mutual expectations, etc.) can be \nused in processing not only formulaic language expressions but a larger set of expressions \nwith a completely transparent meaning. Psycholinguistic literature has studied two \nnoncompositional cases: idiomatic processing, associated with faster reading time [1] and a \nmore positive electric signal in brain activity [2] to transparent phrases, and frequency effects, \ni.e., multiword sequences are usually read faster than comparable sequences of lesser \nfrequency [3,4]. We assume that facilitation effects are not limited to formulaic expressions \nbut also occur when processing highly prototypical and yet compositional phrases. To test this \nhypothesis, we implemented a Self-Paced Reading experiment to compare the Reading \nTimes (RTs) of argument constructions with a different degree in compositionality: idiomatic \nexpressions (ID), compositional highly frequent expressions (HF), and compositional low- \nfrequent expressions (LF). To the best of our knowledge, no previous work had compared \nboth idioms and frequent constructions, except for [5]. Given the previous literature, we \nhypothesized that RTs are longer for compositional sentences than for idiomatic sentences, \nand RTs are longer for infrequent phrases than for frequent ones. \n \nDesign. We selected 48 idiomatic VERB+determinant+NOUN phrases and corresponding \nhigh-frequency and low-frequency bigrams with the same verb. Each stimulus consisted of a \ncontext sentence presented for the participant to read in one instance and a sentence with the \ntarget phrase embedded (Table1), displayed word-by-word using the moving-window SPR \nparadigm [6]. The stimuli were split into three counterbalanced lists randomly initialized at \neach time. The experiment was delivered remotely, and participants were recruited using \nProlific. We collected responses for 90 subjects from the United States and Canada, all self - \nreported L1 speakers of English aged between 18 and 50. \n \nData Analysis. We removed the outliers and examined the RTs of the last word of phrases \nusing linear mixed models. Condition, Age, WordLength, VerbFrequency, and PositionInList \nwere entered in the models as fixed effects; Subject and Item were treated as random effects \nwith a by-subject random slope for BigramFrequency. RTs' difference between ID and HF \nturned out to be not statistically significant (Table2), while it was statistically different between \nID and LF. Changing the reference level with HF condition, there is still a smaller statistical \ndifference between HF and LF. Moreover, we observed that 1) older adults are slower than \nyounger speakers, and 2) RTs at the end of the experiment are faster than at the beginning. \n \nDiscussion. Analysis reveals no difference between processing the figurative meaning of \nidioms and the compositional one of HF; there are facilitation effects in comprehension of both \nexpressions. Even if this observed measure cannot say what is happening at the brain level, \nit opens to a broad discussion about underlying mechanisms in language processing. It may \nsupport the hypothesis that HF expressions are stored as unanalyzed wholes and directly \nretrieved once recognized as idioms, following usage-based models [7,8]. An alternative \nexplanation is the existence of a co-activated network of representations operating together \nwith analogy-based mechanisms leading to sentence meaning construction and working side \nby side with classical compositional ones. RTs for infrequent phrases were significantly \nslower, even if the advantage was relatively small. We presume that information introduced in \ncontext sentences reduces the effort to interpret unpredictable expressions. \n \n[1] Conklin, K., & Schmitt, N. (2008). Formulaic sequences: Are they processed more quickly than \nnonf ormulaic language by native and nonnative speakers?. Applied linguistics, 29(1), 72-89. \n[2] Vespignani, F., Canal, P., Molinaro, N., Fonda, S., & Cacciari, C. (2010). Predictive mechanisms in \nidiom comprehension. Journal of Cognitive Neuroscience, 22(8), 1682-1700. \n[3] Arnon, I., & Snider, N. (2010). More than words: Frequency ef f ects f or multi -word phrases. Journal \nof memory and language, 62(1), 67-82. \n[4] Tremblay, A., Derwing, B., Libben, G., & Westbury, C. (2011). Processing advantages of lexical \nbundles: Evidence f rom self ‐paced reading and sentence recall tasks. Language learning, 61(2), 569- \n613. \n[5] Jolsvai, H., McCauley, S. M., & Christiansen, M. H. (2020). Meaningf ulness beats f requency in \nmultiword chunk processing. Cognitive Science, 44(10). \n[6] Just, M. A., Carpenter, P. A., & Woolley, J. D. (1982). Paradigms and processes in reading \ncomprehension. Journal of experimental psychology: General, 111(2), 228–238. \n[7] Goldberg, A. E. (2006). Constructions at work: The nature of generalization in language. Oxf ord \nUniversity Press on Demand. \n[8] Bybee, J. (2010). Language, usage and cognition. Cambridge University Press. \n[9] Bannard, C. and D. Matthews (2008). Stored word sequences in language learning: The ef f ect of \nf amiliarity on children’s repetition of f our-word combinations. Psychological science 19.3, 241–248.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».