ChatGPT and EFL/ESL Writing: A Systematic Review of Advantages and Challenges
Bibliographic record
Abstract
Artificial Intelligence (AI) has initiated a new era in education with significant potential to revolutionize teaching and learning. ChatGPT has attracted scholars' interest in exploring its beneficial aspects and constraints for enhancing teaching methods and learning experiences. This systematic review aims to investigate the advantages and challenges of ChatGPT for EFL/ESL writing. The data was gathered from three databases, namely Web of Science, Science Direct, and ERIC, between November 2022 and January 2024. A total of 182 publications were gathered using precise keywords, and the 15 most pertinent articles were selected following the PRISMA flowchart. Findings show that ChatGPT can enhance writing efficiency and creativity, improve writing proficiency, and personalize learning experiences. It can also reduce teachers' workload by offering automated evaluation, feedback, and support for revision. The results also show that despite its potential advantages for both instructors and learners, it presents some challenges, such as overreliance, decreased motivation, learning loss, identifying errors at the deep level of the language, offering inconsistent and complex feedback, and issues with academic integrity and originality. The detailed findings and their implications for practice and policy are discussed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".