Assessing the quality and making appropriate use of historical negative control data: A report of the International Workshop on Genotoxicity Testing ( <scp>IWGT</scp> )
Bibliographic record
Abstract
Historical negative control data (HCD) have played an increasingly important role in interpreting the results of genotoxicity tests. In particular, Organisation for Economic Co-operation and Development (OECD) genetic toxicology test guidelines recommend comparing responses produced by exposure to test substances with the distribution of HCD as one of three criteria for evaluating and interpreting study results (referred to herein as "Criterion C"). Because of the potential for inconsistency in how HCD are acquired, maintained, described, and used to interpret genotoxicity testing results, a workgroup of the International Workshops for Genotoxicity Testing was convened to provide recommendations on this crucial topic. The workgroup used example data sets from four in vivo tests, the Pig-a gene mutation assay, the erythrocyte-based micronucleus test, the transgenic rodent gene mutation assay, and the in vivo alkaline comet assay to illustrate how the quality of HCD can be evaluated. In addition, recommendations are offered on appropriate methods for evaluating HCD distributions. Recommendations of the workgroup are: When concurrent negative control data fulfill study acceptability criteria, they represent the most important comparator for judging whether a particular test substance induced a genotoxic effect. HCD can provide useful context for interpreting study results, but this requires supporting evidence that (i) HCD were generated appropriately, and (ii) their quality has been assessed and deemed sufficiently high for this purpose. HCD should be visualized before any study comparisons take place; graph(s) that show the degree to which HCD are stable over time are particularly useful. Qualitative and semi-quantitative assessments of HCD should also be supplemented with quantitative evaluations. Key factors in the assessment of HCD include: (i) the stability of HCD over time, and (ii) the degree to which inter-study variation explains the total variability observed. When animal-to-animal variation is the predominant source of variability, the relationship between responses in the study and an HCD-derived interval or upper bounds value (i.e., OECD Criterion C) can be used with a strong degree of confidence in contextualizing a particular study's results. When inter-study variation is the major source of variability, comparisons between study data and the HCD bounds are less useful, and consequentially, less emphasis should be placed on using HCD to contextualize a particular study's results. The workgroup findings add additional support for the use of HCD for data interpretation; but relative to most current OECD test guidelines, we recommend a more flexible application that takes into consideration HCD quality. The workgroup considered only commonly used in vivo tests, but it anticipates that the same principles will apply to other genotoxicity tests, including many in vitro tests.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".