MétaCan
Menu
Back to cohort
Record W2268831157

Corpus-based Machine Translation: Its Current Development and Perspectives

2015· article· en· W2268831157 on OpenAlexaboutno aff
Zhou Da-jun, Yun Wang

Bibliographic record

VenueInternational Forum of Teaching and Studies · 2015
Typearticle
Languageen
FieldComputer Science
TopicNatural Language Processing Techniques
Canadian institutionsnot available
Fundersnot available
KeywordsMachine translationComputer scienceArtificial intelligenceNatural language processingExample-based machine translationCorpus linguisticsParallel corporaComputer-assisted translationMachine translation software usabilityRule-based machine translationText corpusComputational linguistics
DOInot available

Abstract

fetched live from OpenAlex

A Review of Corpus-based Machine TranslationCorpus and Machine TranslationCorpus is a large-scale database with tremendous collective linguistic information in real use, which is provided for retrieval by computers for research. The first corpus was established in Brown University, U.S.A., in the late 1960s (Zhang & Zhang, 2010, p. 55). Much progress has been made in corpus research and application in the past decades. Current studies of parallel or multi-language corpus can be categorized into three aspects: the first is the alignment technology of parallel language material, with various strategies and approaches provided by scholars and with numbers of programs and tools of alignment parallel or multi-language material; the second is the application of parallel language materials, such as statistic-based machine translation, example-based machine translation, and parallel language dictionary compiling; the third refers to issues of parallel language corpus design and management, and its material collection and coding (Chang, et al., 2003, p. 28).Corpus used in translation is one of the focuses of corpus application research. Machine translation (MT) is a technology to translate from one natural language in character or speech form into another by means of computer programs (Zhao & Liu, 2010, p. 36). MT was initiated in the 1950s and entered into a prosperous period in the end of 1980s, which characterizes practicality of many translation systems in various fields. An English-French translation system TAUMMETEO developed by the University of Montreal, Canada, in 1976, is a typical example, which can provide high-quality translation of weather forecast (Shao, 2010, p. 28). A typical MT system adopts a transfer-based translation strategy, which consists of 3 procedures: 1) analyzing a source language and form representation of the language; 2) transforming representation of the source language into that of the target language; 3) generating the source language translation version from the target language representation (He, 2007, p. 191). Traditional MT has two defects: one is that traditional MT regards words as its basic translation unit; that is, the machine first segments sentences of a source language into words, which are transformed into those of target language, and then those words are connected according to grammatical structure rules of the target language; the other is that traditional MT does not pay much attention to contexts. Peer-to-peer studies of corpus-based MT do attempt to get over the defects of the traditional MT systems and improve efficiencies and accuracy of the MT systems.With the development over more than 50 years, MT systems have performed certain functions in some fields. However, the current systems have not reached the effect of translation as expected. In its earlier period, MT research was conducted from the viewpoint of natural linguistics, thus creating MT systems based on such linguistic rules as lexical rule, syntactic analysis rule, transformation rules, and target language generation rules. As these rules were summarized and developed from experiences of linguists, there exist some deficiencies in the analytic rules. For instance, manual writing of those rules demands a large quantity of workload, and the rules are too subjective to keep consistency (Wang, 2003, p. 33). Since 1989, MT has entered a new stage in which corpus methods are introduced to the rule-based technology, including statistic-based and example-based methods, and the method of turning corpus into linguistic knowledge bank through language material processing, etc. (Feng, 2010, p. 28). The past years have seen the prominent achievement in MT systems.Corpus Used in Machine TranslationThere are three kinds of corpus concerning MT: parallel corpus, multi-language corpus, and comparable corpus. Parallel corpus collects original text of a language and its translating text of another language. …

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Other design · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.856
Threshold uncertainty score0.256

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.090
GPT teacher head0.362
Teacher spread0.272 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designOther design
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations4
Published2015
Admission routes1
Has abstractyes

Explore more

Same venueInternational Forum of Teaching and StudiesSame topicNatural Language Processing TechniquesFrench-language works237,207