Bibliographic record
Abstract
Modern AI systems are trained using sophisticated machine learning algorithms based on large data sets from a variety of sources. However, users that query these systems for information often have little knowledge about how they were trained or what information was used for training. As a result, users may believe the answers they are given to queries even in cases where they would not have trusted the data that was used to train the system. In this paper, we argue that trust in AI systems therefore relies heavily on transparency around the sources and methods used for training. In order to make this point precise, we introduce a model of a source network along with formal belief change operators that indicate how a user's beliefs should change when a trained system provides information. Using this formal framework, we demonstrate that there are cases where an agent can be deceived into believing information provided by an AI system, even if they would not have believed the information if it came directly from the sources used for training. We also show that our formal framework can be used to precisely state desirable properties for AI systems, which will guarantee that the system is only trusted when the underlying sources are trusted. Ethical considerations are discussed, highlighting the problems that occur when systems are allowed to obscure either the algorithms or the training data used.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".