Secondary School EFL Teachers’ Formative Assessment Practices and Their Impact on Learning
Bibliographic record
Abstract
The positive impact of formative assessment (FA) on learning was taken as a conventional wisdom in education for decades, yet the empirical evidence supporting its prospective benefits on learning especially from students’ perspectives remains distinctively lacunary. To fill this gap, this research project aimed at examining the FA practices employed by Tunisian EFL teachers and their impact on students’ learning by putting students’ perspectives at the center of the debate. Semi-structured interviews (n = 5) were addressed to 5 secondary school teachers to examine the FA practices they implemented in their classrooms. Students (n = 100) were administered an internet-based survey to probe their insights as to the impact of these practices on their learning. To cross-check students’ answers to the internet-based survey, comparisons of their test scores between two summative tests (STs) occurring before and after the implementation of the specific FA practices were conducted. Major results showed that EFL teachers referred to providing their students with oral and written feedback, sharing with them the used assessment criteria, and enhancing peer and self-assessment. While students believed that these FA practices are helpful as they enabled them to determine their strengths and weaknesses and to identify what to do to improve their learning, no significant improvement was found in students’ test scores between the two STs. Moreover, students seem to always favor their teachers’ assessment over that of their peers or themselves. These results challenge the entrenched beliefs about FA as the ultimate tool to enhance students’ learning outcomes and open up more venues for further research that land more powerful empirical support for its prospective benefits on learning.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.018 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.004 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".