Bibliographic record
Abstract
Modern AI systems are trained using sophisticated machine learning algorithms based on large data sets from a variety of sources. However, users that query these systems for information often have little knowledge about how they were trained or what information was used for training. As a result, users may believe the answers they are given to queries even in cases where they would not have trusted the data that was used to train the system. In this paper, we argue that trust in AI systems therefore relies heavily on transparency around the sources and methods used for training. In order to make this point precise, we introduce a model of a source network along with formal belief change operators that indicate how a user's beliefs should change when a trained system provides information. Using this formal framework, we demonstrate that there are cases where an agent can be deceived into believing information provided by an AI system, even if they would not have believed the information if it came directly from the sources used for training. We also show that our formal framework can be used to precisely state desirable properties for AI systems, which will guarantee that the system is only trusted when the underlying sources are trusted. Ethical considerations are discussed, highlighting the problems that occur when systems are allowed to obscure either the algorithms or the training data used.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.026 | 0.092 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.005 | 0.019 |
| Scholarly communication | 0.012 | 0.023 |
| Open science | 0.002 | 0.011 |
| Research integrity | 0.006 | 0.009 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".