Bibliographic record
Abstract
The preeminent question of the metaphysics of classification is that of whether the world is itself naturally subdivided into kinds of things. Are kinds out there, so to speak, or are they rather artefacts of convention, existing only insofar as classificatory practices are brought to bear by creatures such as ourselves? In this paper, I examine this question from the point of view of the sciences, and more specifically, from the perspec tive of the most fulsome view of the epistemic credentials of the sciences regarding what's 'out there': scientific realism. As I hope to show, ap proaching the metaphysics of classification from the perspective of scientific realism has important consequences for one's very understanding of the perspective itself. Thus, by considering the nature of kinds from this per spective, I aim to shed light not only on the metaphysics of classification, but also on the nature of realism with respect to scientific knowledge. Scientific realism (simply 'realism', henceforth, unless otherwise indicated) is the view that our best scientific theories are true, or approx imately true, or to put it in terms other than truth, that they describe well, or to some significant degree of success, the ontology of parts of the world. There are explicit caveats built into this coarse definition ('best' theories, 'approximate' truth, 'significant degrees' of success), and I will make no attempt to expound these particular qualifications here. A further clarification of the definition, however, furnishes a central motivation for what follows. Realism is often explicated in terms of three sorts of com mitment: a metaphysical commitment to the existence of a mind-inde pendent reality; a semantic commitment to interpret scientific claims literally (or as it is often put, at face value); and an epistemological commitment to regard these claims as furnishing knowledge of both ob servable and unobservable entities and processes. After the demise of
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.016 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.007 | 0.033 |
| Scholarly communication | 0.008 | 0.012 |
| Open science | 0.002 | 0.008 |
| Research integrity | 0.005 | 0.008 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".