Variability in a multiple translation corpus as evidence for cognitive processes
Notice bibliographique
Résumé
Toury (1995) has proposed two laws (of interference and increasing standardization) to account for translational behavior. Translated language has also been described in terms of distinctive features (e.g., simplification and explicitation; cf. Baker, 1993). Based on these models, the field of Translation Studies has made important descriptive and methodological advances, but it has been suggested that it lags behind in theory development and would benefit from improved interaction with cognitive linguistics and psycholinguistics (De Sutter & Lefer, 2020; Halverson & Kotze, 2021). One notable cognitive linguistic framework is offered by Halverson (2017), whose revised “gravitational pull” (RGP) model aims to explain Toury’s laws in terms of the salience of source and target items and the entrenchment of translation pairs. This talk concerns the French translation of English noun sequences (e.g., bank insurance system). These sequences provide an interesting test case for the RGP model. Psycholinguists have investigated how compound processing may elucidate the structure of the mental lexicon (Libben, 2005; Baayen et al., 2010; Gagné, 2011), and cross-linguistic research can extend this to bilingual cognition. Furthermore, the bare juxtaposition of nouns provides efficient information packing in English but poses several challenges for French translators. At the formal level, the [N+N] structure is much less productive in French, which tends to use prepositional post-modification. For longer sequences, word order reversal and accumulating prepositions can impose a significant cognitive load (e.g., disaster relief program coordinator → coordinateur du programme d'aide en cas de catastrophe). At the semantic level, interpretation may be ambiguous due to (1) polysemy of individual constituents; (2) different plausible syntactic relationships between constituents (e.g., [bank insurance] system vs. bank [insurance system]); and (3) different semantic relationships (e.g., bath water might refer to ‘water for a bath’, ‘in the bath’, or ‘from a bath’). These interpretations, left implicit in English, often need to be explicitated in French. The present study uses data from the Multilingual Student Translation corpus (MUST; Granger & Lefer, 2020), which consists of multiple French student translations for English texts specialized in sustainable finance. Three English source texts with 87 noun sequences give rise to 2974 parallel concordances. This access to multiple translator realizations allows me to examine translation variability, i.e., the number of different solutions for a given source instance. This approach, already championed by Malmkjær (1998), remains underexplored to date (but see Castagnoli, 2020). More specifically, I present results from a multifactorial model of translation variability as a function of sequence length, lexicalization, frequency of use, and structural ambiguity. In addition, I describe qualitatively how solutions vary at three structural levels (main constituents, linking prepositions, and inflections), which can reveal how translators interpret and resolve different kinds of ambiguity. While previous RGP studies have mainly operationalized salience by corpus frequencies and explored the under- or overuse of particular forms, I propose that studying multiple translations provides a complementary approach to elucidate the cognitive and linguistic factors underlying translator decisions. References: Baayen, H.R., Kuperman, V. & Bertram, R. (2010). Frequency effects in compound processing. In S. Scalise & I. Vogel (eds), Cross-Disciplinary Issues in Compounding. Amsterdam: John Benjamins, 257-270. Baker, M. (1993). Corpus linguistics and translation studies: Implications and applications. In Baker, M., Francis, G. & Tognini-Bonelli, E. (eds). Text and Technology: In Honour of John Sinclair. 233-250. Amsterdam & Philadelphia: Benjamins. 233-250. Castagnoli, S. (2020). Translation choices compared: Investigating variation in a learner translation corpus. In Granger, S. & Lefer, M.-A. (Eds.). Translating and Comparing Languages: Corpus-based Insights. Selected Proceedings of the Fifth Using Corpora in Contrastive and Translation Studies Conference. Corpora and Language in Use Proceedings 6. Louvain-la-Neuve: Presses Universitaires de Louvain. 25-44. De Sutter, G., & Lefer, M.-A. (2020). On the need for a new research agenda for corpus-based translation studies: A multi-methodological, multifactorial and interdisciplinary approach. Perspectives, 28(1), 1-23. Gagné, C. L. (2011). Psycholinguistic Perspectives. In Lieber, R., & Štekauer, P. (Eds.). The Oxford handbook of compounding. Oxford University Press. 255-271. Granger, S., & Lefer, M.-A. (2020). The Multilingual Student Translation corpus: a resource for translation teaching and research. Language Resources and Evaluation, 54(4), 1183-1199. Halverson, S. L. & Kotze, H. (2021). Sociocognitive constructs in Translation and Interpreting Studies (TIS): Do we really need concepts like norms and risk when we have a comprehensive usage-based theory of language? In Sandra L. Halverson & Álvaro Marín García, eds. Contesting Epistemologies in Translation and Interpreting Studies. London: Routledge. 51-79. Halverson, S. L. (2017). Gravitational pull in translation: Testing a revised model. In De Sutter, G., Lefer, M. A., & Delaere, I. (eds.). Empirical translation studies: New methodological and theoretical traditions. Walter de Gruyter. 9-46. Libben, G. (2005). Everything is psycholinguistics: Material and methodological considerations in the study of compound processing. Canadian Journal of Linguistics, 50(1-4), 267-283. Malmkjær, K. (1998). Love thy Neighbour: Will Parallel Corpora Endear Linguists to Translators? Meta: Translators' Journal, 43(4), 534-541. Toury, G. (1995). Descriptive Translation Studies – and Beyond. Amsterdam: John Benjamins.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,019 | 0,126 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,005 | 0,006 |
| Études des sciences et des technologies | 0,003 | 0,004 |
| Communication savante | 0,004 | 0,003 |
| Science ouverte | 0,001 | 0,005 |
| Intégrité de la recherche | 0,001 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,004 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».