Bibliographic record
Abstract
Measurement is an indispensable part of physical science as well as of commerce, industry, and daily life. Measuring activities appear unproblematic when performed with familiar instruments such as thermometers and clocks, but a closer examination reveals a host of epistemological questions, including: 1. How is it possible to tell whether an instrument measures the quantity it is intended to? 2. What do claims to measurement accuracy amount to, and how might such claims be justified? 3. When is disagreement among instruments a sign of error, and when does it imply that instruments measure different quantities? Currently, these questions are almost completely ignored by philosophers of science, who view them as methodological concerns to be settled by scientists. This dissertation shows that these questions are not only philosophically worthy, but that their exploration has the potential to challenge fundamental assumptions in philosophy of science, including the distinction between measurement and prediction. The thesis outlines a model-based epistemology of physical measurement and uses it to address the questions above. To measure, I argue, is to estimate the value of a parameter in an idealized model of a physical process. Such estimation involves inference from the final state (‘indication’) of a process to the value range of a parameter (‘outcome’) in light of theoretical and statistical assumptions. Idealizations are necessary preconditions for the possibility of justifying such inferences. Similarly, claims to accuracy, error and quantity individuation can only be adjudicated against the background of an idealized representation of the measurement process. Chapters 1-3 develop this framework and use it to analyze the inferential structure of standardization procedures performed by contemporary standardization bureaus. Standardizing time, for example, is a matter of constructing idealized models of multiple atomic clocks in a way that allows consistent estimates of duration to be inferred from clock indications. Chapter 4 shows that calibration is a special sort of modeling activity, i.e. the activity of constructing and testing models of measurement processes. Contrary to contemporary philosophical views, the accuracy of measurement outcomes is properly evaluated by comparing model predictions to each other, rather than by comparing observations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.016 | 0.034 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.003 | 0.020 |
| Scholarly communication | 0.011 | 0.021 |
| Open science | 0.004 | 0.006 |
| Research integrity | 0.005 | 0.008 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".