xml:lang="en">Lesson study as an approach to facilitate the integration of Gen-AI into EFL curriculum design in higher education
Bibliographic record
Abstract
Purpose This study investigates how English as a Foreign Language (EFL) teachers from higher education develop and refine their curriculum design with Generative Artificial Intelligence (Gen-AI) collaboration during the Lesson Study (LS). Design/methodology/approach Through a qualitative case study approach, we followed six English teachers in their collaborative work with a Gen-AI teaching assistant (Kimi) over a 6-month semester. Data were collected through the recordings of LS cycles, teacher interviews and reflections and documentation of teacher-AI interactions etc. Findings The findings revealed three key aspects of Gen-AI integration in designing EFL curriculum: First, teachers progressively discovered Kimi’s capabilities in lesson planning, material development, and activity design, showing value in generating differentiated learning resources. Second, the teachers developed sophisticated collaboration patterns with the Gen-AI, demonstrating iterative refinement approaches and strategic integration of Gen-AI suggestions throughout the LS cycles. Third, teachers' critical reflections showed evolution in their evaluation and application of Gen-AI contributions, maintaining professional agency while leveraging Gen-AI capabilities effectively. Research limitations/implications This study has several limitations that inform future research directions. Our investigation focused specifically on EFL higher education using a single Gen-AI tool (Kimi), which may limit the generalizability of the findings to other educational contexts and AI platforms. Practical implications These findings suggest that Gen-AI integration through LS can enhance teachers' professional practice while promoting critical engagement with Gen-AI tools. The study provides insights into how Gen-AI can be meaningfully integrated into teacher professional development through collaborative LS approaches. Originality/value The study demonstrates how the LS framework supports balanced AI integration while maintaining teacher agency. In addition, it reveals the process of AI capability discovery and strategic implementation in EFL teaching.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.013 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.023 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".