Reforming Language Teaching through Learner Portfolio Assessment: The Complexity of Practitioners’ Experiences
Bibliographic record
Abstract
This mixed methods study examined practitioners’ experiences with mandatory portfolio-based language assessment (PBLA) as implemented in government-funded language programs for adult newcomers to Canada. Participants’ reports on PBLA impact on language teaching and learning were collected in two phases through online surveys (N1 = 323; N2 = 94) from three participant groups: teachers, PBLA Lead Teachers, and program administrators. Within an overarching framework of complexity theory (CT), the mixed methods survey data were analysed through the lens of the third-generation activity theory (AT) to illuminate potential sources of tensions and contradictions that became prominent after an initial grounded-theory-driven analysis of the dataset. Quantitatively, exploratory factor analysis identified dimensions of practitioner experiences with PBLA. Qualitatively, open-ended comments were analysed through abductive coding cycles leading to a thorough description of PBLA impact on LINC activity systems. Changes between the two data collection phases were traced both through individual factor scores and through quantitization of qualitative data, both of which revealed lack of positive dynamic in practitioner experiences and their appraisal of PBLA impact. The findings provided evidence for (1) re-asserting the key role of individual agency in language teaching and learning; (2) reconceptualizing learner metacognitive awareness as an emergent property of complex adaptive systems; (3) re-centering the organizing role of system activity in achieving system outcomes; (4) problematizing the linearity assumptions behind assessment washback; (5) questioning the use of power in mandatory assessment protocols for presumed improvements in language teaching and learning. Contrary to the expectations of the linear positive washback effect of the mandatory assessment protocol, the complex ecosystems of newcomer language teaching and learning experienced serious unintended consequences, undermining projected benefits. It appears that the linear behaviourist assumptions behind the washback effect underestimated the complexity of (language) teaching and learning in the settlement context. When targeting improvements in education systems, addressing the complexities of teaching and learning process may be a more promising approach than expecting such improvements to be achieved through the washback effect of assessment policies and practices.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.004 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.021 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".