Structured Scaffolding with Digital Graded Audiobooks: Impacts on L2 Listening Comprehension and Learner Perceptions
Bibliographic record
Abstract
This study employed a mixed-methods research design to investigate the effects of structured scaffolding with digital graded audiobooks on the listening comprehension of low-proficiency Thai EFL undergraduates and to explore their perceptions of audiobook use within a structured extensive listening framework. Seventy-four students participated, including thirty-eight in the experimental group and thirty-six in the control group. Research instruments comprised a listening comprehension pretest and posttest, together with semi-structured interviews. The experimental group engaged with twelve Oxford Bookworms digital graded audiobooks integrated into pre-, while-, and post-listening scaffolding activities over twelve weeks, supplemented by autonomous practice through the Oxford Reading Club platform. The control group received conventional textbook-based listening instruction. Quantitative results showed that the experimental group’s mean score improved from 13.03 (SD = 2.17) to 17.68 (SD = 2.76), while the control group increased from 12.69 (SD = 2.59) to 14.86 (SD = 3.39). A significant between-group difference was found at posttest, t(72) = 3.94, p < .001, d = 0.92, indicating a large effect size. Qualitative findings revealed that students perceived digital audiobooks as enhancing comprehension, sound discrimination, accent awareness, vocabulary development, and authentic language use. The study highlights the pedagogical value of digital audiobooks in under-resourced, low-input contexts where access to authentic English and technology is limited. Overall, the results demonstrate that a guided-to-autonomous structured scaffolding model can promote accessible, sustainable, and independent listening development in EFL learning environments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".