MétaCan
Menu
Retour à la cohorte
Enregistrement W7111676231

How Effective Is Data-Driven Learning in Pre-Tertiary Language Education? A Systematic Review

2025· other· en· W7111676231 sur OpenAlexaboutno aff

Notice bibliographique

RevueOSF Preprints (OSF Preprints) · 2025
Typeother
Langueen
Domaine
Thématique
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésSystematic reviewGuidelineConsistency (knowledge bases)Psychological interventionProtocol (science)Plain languagePlan (archaeology)Inclusion (mineral)
DOInon disponible

Résumé

récupéré en direct d'OpenAlex

Drawing on Booth et al. (2016), Cooper (2017) and Higgins et al. (2019), and following the PRISMA 2020 framework, this review synthesises findings from primary studies on DDL in pre-tertiary education published since 2015. It addresses the following questions: 1. Does DDL generate positive language learning outcomes among pre-tertiary learners? 2. What instructional approaches are implemented in DDL practices? 3. What are the reported advantages and disadvantages of using corpora in the language classroom? To ensure consistency and minimise bias, a structured protocol with predefined inclusion criteria was established before the literature search. The search plan followed the PRESS guideline for systematic reviews (McGowan et al., 2016), which prioritises high quality database searching. The PICO framework (Richardson et al., 1995) guided formulation of the review question and informed both the eligibility criteria and the organisation of studies for synthesis. Embedding such frameworks is recognised as a core element of effective review planning (Higgins et al., 2019). References: Booth, A., Sutton, A., & Papaioannou, D. (2016). Systematic approaches to a successful literature review (2nd ed.). SAGE Publications. Cooper, H. (2017). Research synthesis and meta-analysis. SAGE Publications, Inc. Higgins, J. P. T., Thomas, J., Chandler, J., Cumpston, M., Li, T., Page, M. J., & Welch, V. A. (Eds.). (2019). Cochrane handbook for systematic reviews of interventions (2nd ed.). John Wiley & Sons. McGowan, J., Sampson, M., Salzwedel, D. M., Cogo, E., Foerster, V., & Lefebvre, C. (2016). PRESS Peer Review of Electronic Search Strategies: 2015 guideline statement. Journal of Clinical Epidemiology, 75, 40–46. Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., … Moher, D. (2021). The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ, 372, n71. Background Since the early 2000s, multiple systematic reviews and meta-analyses in applied linguistics have examined DDL and consistently report positive effects on language learning, especially for vocabulary and lexicogrammar. Effects are strongest in higher education and with English learners. Hands-on activities are reported more often than hands-off and are linked to better engagement and retention. Classroom uptake remains uneven because of learner, teacher, task, and tool constraints. Challenges include cognitive load from concordance reading, limited teacher training, complex interfaces and queries. Common limitations include small samples, short interventions, inconsistent methods and reporting, over reliance on university participants. Publication and language biases are noted. Common limitations include small samples, short interventions, and over reliance on university participants (Boulton & Cobb, 2017; Boulton & Vyatkina, 2021; Lee et al., 2019; Lusta et al., 2023; Luo, 2025; Pérez-Paredes, 2022; Sun & Mizumoto, 2025; Sun & Park, 2023; Ueno & Takeuchi, 2023). Most syntheses include studies only up to 2016, 2020, or 2022, so an updated review is required to capture recent work and technological developments. Corpus tools have become more accessible, for example new features that help users formulate queries more naturally, which may change how DDL is used and learned from. Evidence on pre-tertiary and lower proficiency learners remains scarce, limiting generalisability beyond universities. A new review will focus on measured learning outcomes, compare delivery modes, and map barriers to implementation in pre-tertiary settings to support wider classroom use of DDL. Information sources The search will target multiple sources to maximise coverage. NUsearch, the University of Nottingham’s library discovery tool, will be used to identify and access databases. Databases will be searched via the ProQuest interface, alongside the Scopus platform. In addition, Taylor and Francis Online, Oxford Academic, and ScienceDirect will be searched to maintain a rigorous search and to capture key journals in language education and applied linguistics. Reference list checking of included studies will be undertaken to identify additional records. Search strategy and terms Boolean operators (AND, OR, NOT) and truncation will be used to build comprehensive and reproducible queries for DDL and corpus use in language classrooms. Database specific strategies, limits, and filters will be provided in Appendix 1. An additional Google Scholar search will use: ‘(data driven learning OR DDL OR corpus based language learning) AND (pre tertiary education OR younger learners) NOT (university OR tertiary OR undergraduate OR graduate)’, and will be limited to 2015 to 2025. Eligibility criteria Studies will be eligible if they: 1. report empirical DDL activities with pre tertiary language learners 2. provide a clear description of the intervention 3. report quantitatively measured L2 learning outcomes, typically via pre and post tests 4. are published in established, peer reviewed journals or edited volumes 5. fall within the 2015 onwards window Exclusions: MA theses, conference outputs and presentations; studies in non-educational contexts; instructional guidelines or purely theoretical work; empirical studies without a clear intervention or measurable outcomes; purely qualitative designs; studies conducted solely in tertiary education; language institute contexts serving adults. Screening and selection All records from databases and supplementary sources will be imported to EndNote for organisation and deduplication. Study selection will follow a three-stage process: title screening, abstract screening against the predefined criteria, and full text screening using the complete eligibility set. Reasons for exclusion at full text will be recorded. Screening and coding will be conducted by a single reviewer, so inter rater reliability checks will not be applicable. A PRISMA 2020 flow diagram will be used to document the process. Data extraction (coding) A standardised data extraction form will capture bibliographic details, context, population and learner profile, baseline data, intervention characteristics, DDL mode, duration, outcome measures, and reported statistics, together with methodological features relevant to risk of bias. The coding categories are as follows: 1 Study ID number 2 Author(s) 3 Title of study 4 Year of publication 5 Type of publication (for example journal article, book chapter) 6 Country of origin 7 Citation 8 Objectives of study (outlined research questions) 9 Methodological design (qualitative, quantitative, or mixed methods) 10 Research design (for example experimental or quasi experimental) 11 Recruitment procedures (evidence of randomisation, group allocation) 12 Duration of intervention 13 Setting (educational institution, level of education, location, classroom or online) 14 Target language and learning context (for example SLA or FLA) 15 Population characteristics (for example sample size, gender, nationality, first language(s)) 16 Learner profile (school year, age range, proficiency level) 17 Baseline characteristics (for example pretest results) 18 Attrition (drop out rate in each group) 19 Type of DDL activities (paper based or computer based with tools specification) 20 Description of the intervention 21 Learning outcome measures (for example pre tests, post tests, delayed tests) 22 Reported statistical analyses (for example effect sizes, Cronbach’s α, SD, and p values) Risk of bias assessment Methodological quality will be appraised using the revised Cochrane’s Risk of Bias assessment tool (Higgins et al., 2019) an adapted MMAT (2018) tailored to educational research. A decision-based approach will categorise studies as high, medium, or low quality. Studies rated low quality or failing essential items will be excluded from synthesis. Synthesis methods A narrative synthesis will be undertaken following guidance from the Centre for Reviews and Dissemination and Booth. First, a descriptive summary and tabulation of included studies will be prepared. Second, effectiveness will be synthesised by summarising statistical findings and learning outcomes across studies. Third, delivery methods will be grouped by degree of mediation. Fourth, barriers to implementation will be identified through an inductive thematic synthesis of discussion and conclusion sections. References: Boulton, A., & Cobb, T. (2017). Corpus use in language learning: A meta-analysis. Language Learning, 67(2), 348–393. https://doi.org/10.1111/lang.12224 Wiley Online Library Boulton, A., & Vyatkina, N. (2021). Thirty years of data-driven learning: Taking stock and charting new directions over time. Language Learning & Technology, 25(3), 66–89. https://doi.org/10.64152/10125/73450 lltjournal.org Hong, Q. N., Pluye, P., Fàbregues, S., Bartlett, G., Boardman, F., Cargo, M., Dagenais, P., Gagnon, M.-P., Griffiths, F., Nicolau, B., O’Cathain, A., Rousseau, M.-C., & Vedel, I. (2018). Mixed Methods Appraisal Tool (MMAT), version 2018. Canadian Intellectual Property Office. http://mixedmethodsappraisaltoolpublic.pbworks.com/ Lee, H., Warschauer, M., & Lee, J. H. (2019). The effects of corpus use on second language vocabulary learning: A multilevel meta-analysis. Applied Linguistics, 40(5), 721–753. https://doi.org/10.1093/applin/amy012 SpringerLink Lusta, A., Demirel, Ö., & Mohammadzadeh, B. (2023). Language corpus and data driven learning (DDL) in language classrooms: A systematic review. Heliyon, 9(12), e22731. https://doi.org/10.1016/j.heliyon.2023.e22731 PubMed Luo, H. (2025). Data-driven learning in second language writing: A systematic review of efficacy, challenges, and futu

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,073
score de la tête « metaresearch » (Gemma)0,194
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Revue systématique · Signal consensuel: Revue systématique
GenreSignal candidat: Synthèse · Signal consensuel: Synthèse
Score de désaccord entre enseignants0,073
Score d'incertitude au seuil0,386

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0730,194
Méta-épidémiologie (sens strict)0,0020,002
Méta-épidémiologie (sens large)0,0130,011
Bibliométrie0,0190,016
Études des sciences et des technologies0,0020,003
Communication savante0,0080,009
Science ouverte0,0030,004
Intégrité de la recherche0,0040,003
Charge utile insuffisante (le modèle a refusé de juger)0,0040,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,010
Tête enseignante GPT0,296
Écart entre enseignants0,286 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeRevue systématique
Domainenon disponible
GenreSynthèse

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2025
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueOSF Preprints (OSF Preprints)Travaux en français237 207