Deploying elements of scoping review methods for adverse outcome pathway development: a space travel case example
Bibliographic record
Abstract
PURPOSE: Health protection agencies require scientific information for evidence-based decision-making and guideline development. However, vetting and collating large quantities of published research to identify relevant high-quality studies is a challenge. One approach to address this issue is the use of adverse outcome pathways (AOPs) that provide a framework to assemble toxicological knowledge into causally linked chains of key events (KEs) across levels of biological organization to culminate in an adverse health outcome of significance to regulatory decision-making. Traditionally, AOPs have been constructed using a narrative review approach where the collection of evidence that supports each pathway is based on prior knowledge of influential studies that can also be supplemented by individually selecting and reviewing relevant references. OBJECTIVES: We aimed to create a protocol for AOP weight of evidence gathering that harnesses elements of both scoping review methods and artificial intelligence (AI) tools to increase transparency while reducing bias and workload of human screeners. METHODS: To develop this protocol, an existing space-health AOP in the workplan of the Organisation for Economic Co-operation and Development (OECD) AOP Programme was used as a case example. To balance the benefits of both scoping review tools and narrative approaches, a study protocol outlining a screening and search strategy was developed, and three reference collection workflows were tested to identify the most efficient method to inform weight of evidence. The workflows differed in their literature search strategies, and combinations of software tools used. RESULTS: Across the three tested workflows, over 59 literature searches were completed, retrieving over 34,000 references of which over 3300 were human reviewed. The most effective of the three methods used a search strategy with searches across each component of the AOP network, SWIFT Review as a pre-filtering software, and DistillerSR to create structured screening and data extraction forms. This methodology effectively retrieved relevant studies while balancing efficiency in data retrieval without compromising transparency, leading to a well-synthesized evidence base to support the AOP. CONCLUSIONS: The workflow is still exploratory in the context of AOP development, and we anticipate adaptations to the protocol with further experience. To further the systematicity, future iterations of the workflow could include structured quality assessment and risk of bias analysis. Overall, the workflow provides a transparent and documented approach to support AOP development, which in turn will support the need for rigorous methods to identify relevant scientific evidence while being practical to allow uptake by the broader community.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.293 | 0.342 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.005 |
| Bibliometrics | 0.016 | 0.017 |
| Science and technology studies | 0.007 | 0.005 |
| Scholarly communication | 0.012 | 0.012 |
| Open science | 0.005 | 0.013 |
| Research integrity | 0.008 | 0.005 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".