MétaCan
Menu
Back to cohort
Record W281329835

A More Valid Alternative to TOEFL

2002· article· en· W281329835 on OpenAlexaboutno aff
Ann E. Roemer

Bibliographic record

VenueCollege and university · 2002
Typearticle
Languageen
FieldComputer Science
TopicEducational Technology and Assessment
Canadian institutionsnot available
Fundersnot available
KeywordsTest of English as a Foreign LanguageTest (biology)Mathematics educationLanguage proficiencyPsychologyMedical educationConstruct (python library)Higher educationForeign languageLanguage assessmentPedagogyComputer sciencePolitical scienceMedicineLaw
DOInot available

Abstract

fetched live from OpenAlex

Abstract The purpose of this article is to describe the TOEFL and the APIEL, and to evaluate both tests on three basic types of validity criteria: content, construct, and criterion-related. These criteria are commonly used by test developers to ensure that the decisions made from a test are as accurate and as fair as possible. Admissions officers want to be confident that the decision they make using a test are the best, based on the most complete information about the applicants. Thousands of international students attend American universities and colleges to reach their educational goals. To do so they, like their American peers, must be evaluated by admissions offices, a task complicated even more by the differences in educational systems worldwide. Proficiency in English is one of the criteria for this process, and it is crucial for success at North American universities. Xu's (1991, p.567) findings strongly suggest that English language proficiency is the single most important factor influencing international graduate students' academic coping (a factor that is just as significant for undergraduate students). As important as it is, assessing an individual's command of English is not a simple procedure. The most widely used instrument is the TOEFL, Test of English as a Foreign Language, required at over 4,200 colleges and universities in the United States and Canada. In 2000-2001, more than half a million examinees registered to take the computer-based TOEFL at test centers from Albania to Zimbabwe (ETS 2001c). Despite its widespread acceptance, the TOEFL may not be an accurate measure of English language proficiency. An alternative to the TOEFL is the APIEL, Advanced Placement in International English Language, which may be a more accurate measure of proficiency in English. TOEFL The TOEFL measures English proficiency of a non-native speaker (ETS 200IC). Presently there are two versions of the test: paper-based and computer-based. As of January zooo, all TOEFL examinations in North America have been computerized, and students have been required to take the TWE, Test of Written English. And since October 2000, the computer-based test has been administered in most test centers abroad. The TOEFL is computer-adaptive, meaning that if the examinees' responses are correct, they will next be presented with more difficult questions, but if their responses are incorrect, the next questions will be of lesser or equal difficulty. There are four sections to the TOEFL. Section i, Listening Comprehension, measures the students' ability to comprehend spoken American English. It includes vocabulary and idioms that are commonly used in spoken language. This section contains two parts, both of which present video clips: short conversations between two speakers and mini-lectures. In the first, the students hear a short dialogue and a question about the dialogue. From the four possible answers on the screen, the students choose the best one to the question they have heard. In the second section, the students hear brief lectures of less than two minutes, after which they answer several questions, spoken one time only. The topics of the conversations and talks are varied, but tend to be academic in nature (ETS 2000). Section 2 of the TOEFL, Structure, measures recognition of formal grammar points in English. It contains two parts. In the first, the students are tested on their ability to choose the correct word or phrase to complete a sentence. Below is an example from an actual TOEFL test (ETS 200Ia). The columbine flower____to nearly all of the United States, can be raised from weed in almost any garden. a. native b. how native is c. how native it is d. is native In the second part, they have to recognize the portion of the sentence that is grammatically incorrect. This section is also multiple choice. …

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.019
metaresearch head score (Gemma)0.102
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.037
Threshold uncertainty score0.123

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0190.102
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0050.004
Science and technology studies0.0020.003
Scholarly communication0.0060.007
Open science0.0030.005
Research integrity0.0030.003
Insufficient payload (model declined to judge)0.0370.008

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.024
GPT teacher head0.249
Teacher spread0.225 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations12
Published2002
Admission routes1
Has abstractyes

Explore more

Same venueCollege and universitySame topicEducational Technology and AssessmentFrench-language works237,207