Assessment for learning in a chinese university context: a mixed methods case study on english as a foreign language speaking ability
Bibliographic record
Abstract
This study investigates the effectiveness of Assessment for Learning (AFL) in improving oral English skills and explores students' and teachers' perceptions of AFL.The study took place at a university in China and involved both students and teachers of English at the institution.Chinese university level students were reported to be facing difficulties in their oral skills learning and were not satisfied with the oral English instruction they were receiving because it is related too much to large-scale tests administered in China (He, 1999;Liao & Qin, 2000;Wen, 2001).Classroom-based assessment, known as the alternative assessment approach, has attracted increased interest from researchers since the end of the last century (Genesee & Upshur, 1996;Gipps 1999; Shepard, 2000; Turner, in press).One approach to classroom-based assessment, Assessment for Learning (AFL), has proved a significant influence on language performance by encouraging learners' participation, identifying learners' weaknesses, providing instructors with useful feedback for learners' further development, and turning learners into autonomous learners (Black & Wiliam,
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.003 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".