Reforming Language Teaching through Learner Portfolio Assessment: The Complexity of Practitioners’ Experiences
Bibliographic record
Abstract
This mixed methods study examined practitioners’ experiences with mandatory portfolio-based language assessment (PBLA) as implemented in government-funded language programs for adult newcomers to Canada. Participants’ reports on PBLA impact on language teaching and learning were collected in two phases through online surveys (N1 = 323; N2 = 94) from three participant groups: teachers, PBLA Lead Teachers, and program administrators. Within an overarching framework of complexity theory (CT), the mixed methods survey data were analysed through the lens of the third-generation activity theory (AT) to illuminate potential sources of tensions and contradictions that became prominent after an initial grounded-theory-driven analysis of the dataset. Quantitatively, exploratory factor analysis identified dimensions of practitioner experiences with PBLA. Qualitatively, open-ended comments were analysed through abductive coding cycles leading to a thorough description of PBLA impact on LINC activity systems. Changes between the two data collection phases were traced both through individual factor scores and through quantitization of qualitative data, both of which revealed lack of positive dynamic in practitioner experiences and their appraisal of PBLA impact. The findings provided evidence for (1) re-asserting the key role of individual agency in language teaching and learning; (2) reconceptualizing learner metacognitive awareness as an emergent property of complex adaptive systems; (3) re-centering the organizing role of system activity in achieving system outcomes; (4) problematizing the linearity assumptions behind assessment washback; (5) questioning the use of power in mandatory assessment protocols for presumed improvements in language teaching and learning. Contrary to the expectations of the linear positive washback effect of the mandatory assessment protocol, the complex ecosystems of newcomer language teaching and learning experienced serious unintended consequences, undermining projected benefits. It appears that the linear behaviourist assumptions behind the washback effect underestimated the complexity of (language) teaching and learning in the settlement context. When targeting improvements in education systems, addressing the complexities of teaching and learning process may be a more promising approach than expecting such improvements to be achieved through the washback effect of assessment policies and practices.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.035 | 0.053 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.006 | 0.007 |
| Scholarly communication | 0.009 | 0.008 |
| Open science | 0.002 | 0.010 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".