MétaCan
Menu
← Retour à la cohorte
Enregistrement W1936582270

Improving post-editing and automatic translation by the creation of phraseological databases: an experiment

2014· article· en· W1936582270 sur OpenAlexaboutno aff
Jean-Pierre Colson

Notice bibliographique

RevueDigital Access to Libraries (Université catholique de Louvain (UCL), l'Université de Namur (UNamur) and the Université Saint-Louis (USL-B)) · 2014
Typearticle
Langueen
DomaineComputer Science
ThématiqueNatural Language Processing Techniques
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésPhraseologyLinguisticsContext (archaeology)Computer scienceLexiconMachine translationCorpus linguisticsNatural language processingArtificial intelligenceComputational linguisticsLexical databaseEncyclopediaWordNetHistoryPhilosophy
DOInon disponible

Résumé

récupéré en direct d'OpenAlex

In spite of the success of phraseology across a range of linguistic disciplines such as corpus linguistics, discourse analysis or semantics, it may come as a surprise that the notion is hardly mentioned in Translation Studies. Delisle (2003), for instance, treats set phrases as part of the lexicon. They are also most conspicuously absent from the major reference work in the field, the Routledge Encyclopedia of Translation Studies (Baker and Saldanha 2011). The same holds true of collocations. As a matter of fact, the interest for phraseology in translation studies came mainly from the European Society for Phraseology (Europhras) and from corpus linguistics (e.g. Teubert 2002). Computational linguistics, in its turn, is showing a growing interest for matters involving translation and collocations in the broad sense. It is now generally recognised that phraseology poses a serious problem to machine translation (MT), because it involves a higher semantic level that cannot be grasped by processing the individual words. Multi-word units (including also lexical bundles, Biber et al. 2004) have indeed been called a pain in the neck for NLP (Sag et al. 2001). Recent findings from studies devoted to the performance of MT with regard to phraseology (Monti, Mitkov, Corpas Pastor and Seretan, eds. 2013) suggest that traditional, syntactically based systems obtain lower scores than statistically based systems such as Google Translate. However, Google Translate yields erroneous results for phraseology in at least 40 percent of the cases, which may easily be confirmed by typing randomly chosen collocations in context, especially if they are partly fixed or if the set phrases are less common. I will try to show that a major stumbling block for post-editing or MT remains the very incomplete listing of all set phrases by dictionaries and databases, even for the world's most documented language, English. This obviously has to do with the rather poor results obtained by the automatic extraction of collocations, after no less than 50 years of research (Gries 2013). I will also propose a tentative step in the direction of a better automatic extraction of phraseology in the broad sense, based on the well-known Firthian principle that You shall know a word by the company it keeps (Firth 1957), and on its implementation in terms of metric clusters, a statistical technique derived from IR (Information Retrieval, Baeza-Yates and Ribeiro-Neto 1999). Achieving an appropriate balance between the principles of raw frequency, recurrent frequency and statistical co-occurrence may also be a key to success in future automatic extraction of collocations. The first results yielded by this method are promising, and may already be profitable to post-editing and automatic translation. In the present phase of the experiment, an SQL database of about 400,000 candidate collocations and lexical bundles has be constituted for English by means of the algorithm mentioned above. A demonstration will be shown of a web application enabling users to input a text and discover (part of) its phraseology within a matter of a few seconds. References Baeza-Yates, R. & B. Ribeiro-Neto (1999). Modern Information Retrieval. New York: ACM Press, Addison Wesley. Baker, M. & G. Saldanha (eds.) (2011). Routledge Encyclopedia of Translation Studies. New York: Routledge. Biber, D., Conrad, S. & V. Cortes (2004). “If you look at… Lexical Bundles in University Teaching and Textbooks.” Applied Linguistics 25(3): 371-405. Colson, J.-P. (2007). The World Wide Web as a corpus for set phrases. In: H. Burger, D. Dobrovol’skij, P. Kühn & N. R. Norrick (eds.), Phraseologie / Phraseology. Ein internationales Handbuch der zeitgenössischen Forschung / An International Handbook of Contemporary Research. Volume 2. Berlin, New York: Walter de Gruyter, p. 1071-1077. Colson, J.-P. (2008). Cross-linguistic phraseological studies: An overview. In: Granger, S. & F. Meunier (eds.), Phraseology. An interdisciplinary perspective. John Benjamins, Amsterdam / Philadelphia, p. 191-206. Colson J.-P. (2010a). The Contribution of Web-based Corpus Linguistics to a Global Theory of Phraseology. In: Ptashnyk, S., Hallsteindóttir, E. & N. Bubenhofer (eds.), Corpora, Web and Databases. Computer-Based Methods in Modern Phraselogy and Lexicography. Hohengehren, Schneider Verlag, p. 23-35. Colson, J.-P. (2010b). Automatic extraction of collocations: a new Web-based method. In: S. Bolasco, S., Chiari, I. & L. Giuliano, Proceedings of JADT 2010, Statistical Analysis of Textual Data, Sapienza University of Rome, 9-11 June 2010. Milan, LED Edizioni, p. 397-408. Colson, J.-P. (2012). A new statistical classification of set phrases. In : Pamies, A., Pazos Bretaña, J.M. & L. Luque Nadal (eds.), Phraseology and Discourse : Cross Linguistic and Corpus-based Approaches. Phraseologie und Parömiologie, Band 29. Hohengehren, Schneider Verlag, p. 377-385. Delisle, J. (2003). La traduction raisonnée. Ottawa: Presses de l’Université d’Ottawa. Firth, J. R. (1957). A synopsis of linguistic theory 1930–1955. In F. Palmer (Ed.), Selected Papers of J. R. Firth 1952–1959. London: Longman, p. 168–205. Gries, S. (2013). 50-something years of work on collocations. What is or should be next … International Journal of Corpus Linguistics, 18, p. 137-165. Monti, J., Mitkov, R., Corpas Pastor, G. & V. Seretan (eds) (2013). Workshop Proceedings: Multi-word units in machine translation and translation technologies, Nice 14th Machine Translation Summit. Sag, I., Baldwin, T., Bond, F., Copestake, A. & D. Flickinger (2002). “Multiword expressions: A pain in the neck for NLP”. In: Proccedings of Computational Linguistics and Intelligent Text Processing (CICLing-2002), Lecture Notes in Computer Science, 2276, 1-15. Teubert, W. (2002). “The role of parallel corpora in translation and multilingual lexicography”. In: Altenberg, B. & S. Granger (eds.) Lexis in Contrast: Corpus-based approaches, 189–214.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,004
score de la tête « metaresearch » (Gemma)0,016
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Expérimental (laboratoire) · Signal consensuel: aucune
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,008
Score d'incertitude au seuil0,025

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0040,016
Méta-épidémiologie (sens strict)0,0010,000
Méta-épidémiologie (sens large)0,0020,001
Bibliométrie0,0010,002
Études des sciences et des technologies0,0010,001
Communication savante0,0020,004
Science ouverte0,0020,002
Intégrité de la recherche0,0020,002
Charge utile insuffisante (le modèle a refusé de juger)0,0080,007

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,011
Tête enseignante GPT0,225
Écart entre enseignants0,214 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeExpérimental (laboratoire)
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2014
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueDigital Access to Libraries (Université catholique de Louvain (UCL), l'Université de Namur (UNamur) and the Université Saint-Louis (USL-B))→Même sujetNatural Language Processing Techniques→Travaux en français237 207→