Factors Impacting Recognition of False Collocations by Speakers of English as L1 and L2
Bibliographic record
Abstract
Currently there is a general uncertainty about what makes collocations (i.e., fixedword combinations with specific, not easily interpreted relations between theircomponents) hard for ESL learners to master, and about how to improve collocationrecognition and learning process. This study explored and designed acomparative classification of external factors that affect adult English as L1 andL2 speakers’ recognition of false (non-native-like) collocations. At Stage 1, thereading-comprehension test and interview were administered to 5 participants:emergent bilinguals (another language/English), an advanced bilingual (Russian/English), and a monolingual speaker of English. The results provided insights intothe strategies and criteria participants employed to identify false collocations. AtStage 2, more than 90 speakers of English as L1 and L2 took a reading-comprehensiontest and posttest survey that gathered information on the differences andsimilarities in the participants’ collocation recognition. The results suggested thatcertain factors positively correlated with recognition of false collocations: Englishas a predominant language of communication, the length of residence, vocabularylearningstrategies, and the focus of attention on sentence structure and the formand meaning of word combinations. The implications concern potential focusareas for pedagogical intervention in the ESL classroom.Une incertitude généralisée plane quant à la raison pour laquelle la maitrise desexpressions figées (c.-à-d., des combinaisons de mots unis par des liens spécifiqueset difficiles à interpréter) est problématique pour les apprenants d’ALS, et quantaux moyens d’améliorer la reconnaissance et l’apprentissage des expressionsfigées. Nous avons conçu et ensuite exploré une classification comparative desfacteurs externes qui affectent la reconnaissance, par des adultes locuteurs natifsd’anglais et des adultes dont l’anglais est la langue seconde, de fausses expressionsfigées. La première étape a consisté en un test de compréhension à l’écrit et uneentrevue avec 5 participants: des bilingues émergents (autre langue/anglais), unbilingue avancé (russe/anglais) et un locuteur unilingue d’anglais. Les résultatsoffrent des informations sur les stratégies et les critères qu’emploient les participantspour identifier de fausses expressions figées. Pendant la deuxième étape,plus de 90 locuteurs natifs d’anglais et de locuteurs dont l’anglais est la langueseconde ont complété un examen de compréhension de lecture et une enquête desuivi portant sur les différences et les similarités dans la reconnaissance par lesparticipants des expressions figées. Les résultats laissent entrevoir une corrélationpositive entre certains facteurs et la reconnaissance de fausses expressions figées:le fait d’avoir l’anglais comme langue dominante, la durée du séjour, les stratégies pour apprendre le vocabulaire, et l’attention portée à la structure des phrases ainsi qu’à la forme et au sens des combinaisons de mots. Cette recherche a des retombées quant aux domaines potentiels sur lesquels pourraient porter les interventions pédagogiques dans les cours d’anglais langue seconde.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.162 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".