Variability in a multiple translation corpus as evidence for cognitive processes
Bibliographic record
Abstract
Toury (1995) has proposed two laws (of interference and increasing standardization) to account for translational behavior. Translated language has also been described in terms of distinctive features (e.g., simplification and explicitation; cf. Baker, 1993). Based on these models, the field of Translation Studies has made important descriptive and methodological advances, but it has been suggested that it lags behind in theory development and would benefit from improved interaction with cognitive linguistics and psycholinguistics (De Sutter & Lefer, 2020; Halverson & Kotze, 2021). One notable cognitive linguistic framework is offered by Halverson (2017), whose revised “gravitational pull” (RGP) model aims to explain Toury’s laws in terms of the salience of source and target items and the entrenchment of translation pairs. This talk concerns the French translation of English noun sequences (e.g., bank insurance system). These sequences provide an interesting test case for the RGP model. Psycholinguists have investigated how compound processing may elucidate the structure of the mental lexicon (Libben, 2005; Baayen et al., 2010; Gagné, 2011), and cross-linguistic research can extend this to bilingual cognition. Furthermore, the bare juxtaposition of nouns provides efficient information packing in English but poses several challenges for French translators. At the formal level, the [N+N] structure is much less productive in French, which tends to use prepositional post-modification. For longer sequences, word order reversal and accumulating prepositions can impose a significant cognitive load (e.g., disaster relief program coordinator → coordinateur du programme d'aide en cas de catastrophe). At the semantic level, interpretation may be ambiguous due to (1) polysemy of individual constituents; (2) different plausible syntactic relationships between constituents (e.g., [bank insurance] system vs. bank [insurance system]); and (3) different semantic relationships (e.g., bath water might refer to ‘water for a bath’, ‘in the bath’, or ‘from a bath’). These interpretations, left implicit in English, often need to be explicitated in French. The present study uses data from the Multilingual Student Translation corpus (MUST; Granger & Lefer, 2020), which consists of multiple French student translations for English texts specialized in sustainable finance. Three English source texts with 87 noun sequences give rise to 2974 parallel concordances. This access to multiple translator realizations allows me to examine translation variability, i.e., the number of different solutions for a given source instance. This approach, already championed by Malmkjær (1998), remains underexplored to date (but see Castagnoli, 2020). More specifically, I present results from a multifactorial model of translation variability as a function of sequence length, lexicalization, frequency of use, and structural ambiguity. In addition, I describe qualitatively how solutions vary at three structural levels (main constituents, linking prepositions, and inflections), which can reveal how translators interpret and resolve different kinds of ambiguity. While previous RGP studies have mainly operationalized salience by corpus frequencies and explored the under- or overuse of particular forms, I propose that studying multiple translations provides a complementary approach to elucidate the cognitive and linguistic factors underlying translator decisions. References: Baayen, H.R., Kuperman, V. & Bertram, R. (2010). Frequency effects in compound processing. In S. Scalise & I. Vogel (eds), Cross-Disciplinary Issues in Compounding. Amsterdam: John Benjamins, 257-270. Baker, M. (1993). Corpus linguistics and translation studies: Implications and applications. In Baker, M., Francis, G. & Tognini-Bonelli, E. (eds). Text and Technology: In Honour of John Sinclair. 233-250. Amsterdam & Philadelphia: Benjamins. 233-250. Castagnoli, S. (2020). Translation choices compared: Investigating variation in a learner translation corpus. In Granger, S. & Lefer, M.-A. (Eds.). Translating and Comparing Languages: Corpus-based Insights. Selected Proceedings of the Fifth Using Corpora in Contrastive and Translation Studies Conference. Corpora and Language in Use Proceedings 6. Louvain-la-Neuve: Presses Universitaires de Louvain. 25-44. De Sutter, G., & Lefer, M.-A. (2020). On the need for a new research agenda for corpus-based translation studies: A multi-methodological, multifactorial and interdisciplinary approach. Perspectives, 28(1), 1-23. Gagné, C. L. (2011). Psycholinguistic Perspectives. In Lieber, R., & Štekauer, P. (Eds.). The Oxford handbook of compounding. Oxford University Press. 255-271. Granger, S., & Lefer, M.-A. (2020). The Multilingual Student Translation corpus: a resource for translation teaching and research. Language Resources and Evaluation, 54(4), 1183-1199. Halverson, S. L. & Kotze, H. (2021). Sociocognitive constructs in Translation and Interpreting Studies (TIS): Do we really need concepts like norms and risk when we have a comprehensive usage-based theory of language? In Sandra L. Halverson & Álvaro Marín García, eds. Contesting Epistemologies in Translation and Interpreting Studies. London: Routledge. 51-79. Halverson, S. L. (2017). Gravitational pull in translation: Testing a revised model. In De Sutter, G., Lefer, M. A., & Delaere, I. (eds.). Empirical translation studies: New methodological and theoretical traditions. Walter de Gruyter. 9-46. Libben, G. (2005). Everything is psycholinguistics: Material and methodological considerations in the study of compound processing. Canadian Journal of Linguistics, 50(1-4), 267-283. Malmkjær, K. (1998). Love thy Neighbour: Will Parallel Corpora Endear Linguists to Translators? Meta: Translators' Journal, 43(4), 534-541. Toury, G. (1995). Descriptive Translation Studies – and Beyond. Amsterdam: John Benjamins.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.019 | 0.126 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.006 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.001 | 0.005 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".