Bibliographic record
Abstract
A family of familiar linguistic tests purport to help identify when a term is ambiguous. These tests are philosophically important: a familiar philosophical strategy is to claim that some phenomenon is disunified and its accompanying term is ambiguous. The tests have been used to evaluate disunification proposals about causation, pain, and knowledge, among others. These ambiguity tests, however, have come under fire. It has been alleged that the tests fail for polysemy: a common type of ambiguity, and one that is at issue in philosophically interesting cases. Furthermore, the objection that the tests fail for polysemy is often taken to be an undeniable bit of linguistic data. We argue that this is mistaken. The objection implicitly relies on controversial assumptions about how to account for copredicational sentences, in which a single argument is ascribed prima facie incompatible properties. Furthermore, on several viable theories of copredication, the objection fails. However, our discussion also reveals that even if ambiguity tests are preserved, they may be significantly harder to execute than previously thought.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".