ESL/EFL instructors' practices for writing assessment: specific purposes or general purposes?
Bibliographic record
Abstract
A fundamental difference emerged between specific and general purposes for language assessment in the process of my interviewing 48 highly experienced instructors of ESL/EFL composition about their usual practices for writing assessment in courses in universities or immigrant settlement programs. The instructors worked in situations where English is either the majority language (Australia, Canada, New Zealand) or an international language (Hong Kong, Japan, Thailand). Although the instructors tended to conceptualize ESL/EFL writing instruction in common ways overall, I was surprised to find how their conceptualizations of student assessment varied depending on whether the courses they taught were defined in reference to general or specific purposes for learning English. Conceptualizing ESL/EFL writing for specific purposes (e.g., in reference to particular academic disciplines or employment domains) provided clear rationales for selecting tasks for assessment and specifying standards for achievement; but these situations tended to use limited forms of assessment, based on limited criteria for student achievement. Conceptualizing ESL/EFL writing for general purposes, either for academic studies or settlement in an English-dominant country, was associated with varied methods and broad-based criteria for assessing achievement, focused on individual learners’ development, but realized in differing ways by different instructors.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.056 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".