MétaCan
Menu
Back to cohort
Record W3205076335 · doi:10.1080/02687038.2021.1897081

Utilising a systematic review-based approach to create a database of individual participant data for meta- and network meta-analyses: the RELEASE database of aphasia after stroke

2021· article· en· W3205076335 on OpenAlexaff
Louise R. Williams, Myzoon Ali, Kathryn VandenBerg, Linda Williams, Masahiro Abo, Frank Becker, Audrey Bowen, Caitlin Brandenburg, Caterina Breitenstein, Stefanie Bruehl, David A. Copland, Tamara Cranfill, Marie di Pietro-Bachmann, Pam Enderby, Joanne Fillingham, Federica Galli, Marialuisa Gandolfi, Bertrand Glize, Erin Godecke, Neil Hawkins, Katerina Hilari, Jacqueline Hinckley, Simon Horton, David Howard, Petra Jaecks, Elizabeth Jefferies, Luís M. T. Jesus, Maria Kambanaros, Eun Kyoung Kang, Eman M. Khedr, Anthony Pak‐Hin Kong, Tarja Kukkonen, Marina Laganaro, Matthew A. Lambon Ralph, Ann Charlotte Laska, Béatrice Leemann, Alexander Leff, Roxele Ribeiro Lima, Antje Lorenz, Brian MacWhinney, Rebecca Shisler Marshall, Flavia Mattioli, İlknur Maviş, Marcus Meinzer, Reza Nilipour, Enrique Noé, Nam‐Jong Paik, Rebecca Palmer, Ilias Papathanasiou, Brígida Patrício, Isabel Pavão Martins, Cathy J. Price, Tatjana Prizl Jakovac, Elizabeth Rochon, Miranda L. Rose, Charlotte Rosso, Ilona Rubi‐Fessen, Marina B. Ruiter, Claerwen Snell, Benjamin Stahl, Jerzy P. Szaflarski, Shirley Thomas, Mieke van de Sandt‐Koenderman, Ineke van der Meulen, Evy Visch‐Brink, Linda Worrall, Heather Harris Wright, Marian Brady

Bibliographic record

VenueAphasiology · 2021
Typearticle
Languageen
FieldMedicine
TopicAcute Ischemic Stroke Management
Canadian institutionsToronto Rehabilitation InstituteUniversity of Toronto
FundersQatar National LibraryTavistock Trust for AphasiaHealth Services and Delivery Research ProgrammeMedical Research CouncilNational Institute for Health and Care Research
KeywordsAphasiaDatabasePopulationData collectionStroke (engine)Computer scienceMedicineStatisticsPsychiatry

Abstract

fetched live from OpenAlex

Background Collation of aphasia research data across settings, countries and study designs using big data principles will support analyses across different language modalities, levels of impairment, and therapy interventions in this heterogeneous population. Big data approaches in aphasia research may support vital analyses, which are unachievable within individual trial datasets. However, we lack insight into the requirements for a systematically created database, the feasibility and challenges and potential utility of the type of data collated.Aim To report the development, preparation and establishment of an internationally agreed aphasia after stroke research database of individual participant data (IPD) to facilitate planned aphasia research analyses.Methods Data were collated by systematically identifying existing, eligible studies in any language (≥10 IPD, data on time since stroke, and language performance) and included sourcing from relevant aphasia research networks. We invited electronic contributions and also extracted IPD from the public domain. Data were assessed for completeness, validity of value-ranges within variables, and described according to pre-defined categories of demographic data, therapy descriptions, and language domain measurements. We cleaned, clarified, imputed and standardised relevant data in collaboration with the original study investigators. We presented participant, language, stroke, and therapy data characteristics of the final database using summary statistics.Results From 5256 screened records, 698 datasets were potentially eligible for inclusion; 174 datasets (5928 IPD) from 28 countries were included, 47/174 RCT datasets (1778 IPD) and 91/174 (2834 IPD) included a speech and language therapy (SLT) intervention. Participants’ median age was 63 years (interquartile range [53, 72]), 3407 (61.4%) were male and median recruitment time was 321 days (IQR 30, 1156) after stroke. IPD were available for aphasia severity or ability overall (n = 2699; 80 datasets), naming (n = 2886; 75 datasets), auditory comprehension (n = 2750; 71 datasets), functional communication (n = 1591; 29 datasets), reading (n = 770; 12 datasets) and writing (n = 724; 13 datasets). Information on SLT interventions were described by theoretical approach, therapy target, mode of delivery, setting and provider. Therapy regimen was described according to intensity (1882 IPD; 60 datasets), frequency (2057 IPD; 66 datasets), duration (1960 IPD; 64 datasets) and dosage (1978 IPD; 62 datasets).Discussion Our international IPD archive demonstrates the application of big data principles in the context of aphasia research; our rigorous methodology for data acquisition and cleaning can serve as a template for the establishment of similar databases in other research areas.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.165
metaresearch head score (Gemma)0.459
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.835
Threshold uncertainty score0.874

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1650.459
Meta-epidemiology (narrow)0.0020.003
Meta-epidemiology (broad)0.0110.011
Bibliometrics0.0330.034
Science and technology studies0.0020.001
Scholarly communication0.0100.005
Open science0.0060.010
Research integrity0.0040.003
Insufficient payload (model declined to judge)0.0190.003

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.372
GPT teacher head0.409
Teacher spread0.037 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
DomainMethods
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations9
Published2021
Admission routes1
Has abstractyes

Explore more

Same venueAphasiologySame topicAcute Ischemic Stroke ManagementFrench-language works237,207