Quantitative weight of evidence assessment of higher-tier studies on the toxicity and risks of neonicotinoid insecticides in honeybees 1: Methods
Bibliographic record
Abstract
A quantitative weight of evidence (QWoE) methodology was developed and used to assess many higher-tier studies on the effects of three neonicotinoid insecticides: clothianidin (CTD), imidacloprid (IMI), and thiamethoxam (TMX) on honeybees. A general problem formulation, a conceptual model for exposures of honeybees, and an analysis plan were developed. A QWoE methodology was used to characterize the quality of the available studies from the literature and unpublished reports of studies conducted by or for the registrants. These higher-tier studies focused on the exposures of honeybees to neonicotinoids via several matrices as measured in the field as well as the effects in experimentally controlled field studies. Reports provided by Bayer Crop Protection and Syngenta Crop Protection and papers from the open literature were assessed in detail, using predefined criteria for quality and relevance to develop scores (on a relative scale of 0-4) to separate the higher-quality from lower-quality studies and those relevant from less-relevant results. The scores from the QWoEs were summarized graphically to illustrate the overall quality of the studies and their relevance. Through mean and standard errors, this method provided graphical and numerical indications of the quality and relevance of the responses observed in the studies and the uncertainty associated with these two metrics. All analyses were conducted transparently and the derivations of the scores were fully documented. The results of these analyses are presented in three companion papers and the QWoE analyses for each insecticide are presented in detailed supplemental information (SI) in these papers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".