MétaCan
Menu
← Back to cohort
Record W7008486421

CESTA Evaluation Package

2007· other· en· W7008486421 on OpenAlexaboutno aff

Bibliographic record

VenueAmericanae (AECID Library) · 2007
Typeother
Languageen
FieldComputer Science
TopicNatural Language Processing Techniques
Canadian institutionsnot available
Fundersnot available
KeywordsTerminologyArabicChristian ministrySection (typography)Machine translationDomain (mathematical analysis)
DOInot available

Abstract

fetched live from OpenAlex

The CESTA Evaluation Package was produced within the French national project CESTA (Evaluation of MT systems), as part of the Technolangue programme funded by the French Ministry of Research and New Technologies (MRNT). The CESTA project enabled to carry out a campaign for the evaluation of machine translation systems with English and Arabic texts translated into French. This package includes the material that was used for the CESTA evaluation campaign. It includes resources, protocols, scoring tools, results of the campaign, etc., that were used or produced during the campaign. The aim of these evaluation packages is to enable external players to evaluate their own system. The campaign is distributed over two actions: 1)Evaluation on a restrictive vocabulary: an evaluation protocol was introduced and was dedicated to two translation directions: English into French and Arabic into French.2)Evaluation on a specialised domain (evaluation after terminology enrichment): it consists in observing the impact of the systems adaptation to the specialised domain.The CESTA evaluation package contains the following data and tools:1)Test run data:-English-French parallel corpus: 21,590 English words and 23,554 French words extracted from the Official Journal of the European Communities, 1993, Written Questions section of the European Parliament, from the MLCC corpus (catalogue ref. ELRA-W0023). -Arabic-French parallel corpus: 15,603 Arabic words and 18,257 French words extracted from Le Monde Diplomatique 2002 (catalogue ref. ELRA-W0036).2)First campaign data:-English-French parallel corpus: test corpus of 20,658 English words and 22,774 French words extracted from the Official Journal of the European Communities, 1993, Written Questions section of the European Parliament, from the MLCC corpus (catalogue ref. ELRA-W0023). Four translations in French are available.-Arabic-French parallel corpus: test corpus of 23,763 Arabic words and 28,664 French words extracted from Le Monde Diplomatique 2002 and 2003 (catalogue réf. ELRA-W0036). Four translations in French are available.3)Second campaign data:-English-French parallel corpus: adaptation corpus of 19,383 English words and 22,741 French words, extracted from the Santé Canada website. Translation in French is available.-Arabic-French parallel corpus: adaptation corpus of 19,560 Arabic words and 22,533 French words extracted from the UNICEF, WHO and FHI websites. Translation in French is available.-English-French parallel corpus: test corpus of 18,880 English words and 23,411 French words, extracted from the Santé Canada website. Four translations in French are available.-Arabic-French parallel corpus: test corpus of 17,305 Arabic words and 20,885 French words extracted from the UNICEF, WHO and FHI websites. Four translations in French are available.4)Anonymised submissions of systems and human judgments with adequacy and fluency annotations.5)French corpus of 13,000 words with adequacy and fluency tags.6)Evaluation infrastructure for human judgments and for automatic evaluation.7)Project documentation and publications.A description of the project is available at the following address:http://www.technolangue.net/article.php3?id_article=199 (in French language)

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.016
metaresearch head score (Gemma)0.057
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Other · Consensus signal: none
Teacher disagreement score0.291
Threshold uncertainty score0.975

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0160.057
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0020.003
Bibliometrics0.0060.005
Science and technology studies0.0020.001
Scholarly communication0.0070.005
Open science0.0040.004
Research integrity0.0020.003
Insufficient payload (model declined to judge)0.2910.245

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.012
GPT teacher head0.274
Teacher spread0.262 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreOther

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2007
Admission routes1
Has abstractyes

Explore more

Same venueAmericanae (AECID Library)→Same topicNatural Language Processing Techniques→French-language works237,207→