Bibliographic record
Abstract
Summary form only given. Language is primarily a physical, more particularly a biological phenomenon. To say that it is primarily so is to say that that is how, in the first instance, it presents itself to observation. It is curious then that theoreticians of language treat it as though it were primarily semantic or syntactic or a fusion of the two, and as though our implicit understanding of semantics and the syntax regulates both our language production and our language comprehension. The brain is both a repository of semantic and syntactic constraints, and is the instrument by which we draw upon these accounts for the hard currency of linguistic exchange. With this view comes a division of the vocables of language into those that carry semantic content (lexical vocabulary) and those that mark syntactic form (functional and vocabulary). Logical theory of the past 150 years has been understood by many as a purified abstraction of linguistic forms. So it is not surprising that the logical vocabulary of natural language has been understood in the reflected light of that formal science. Those internal transactions in which logical vocables essentially figure, the transactions that we think of as reasonings, are seen by many as constrained by those laws of thought that logic was thought to codify. Of course no vocabulary can be entirely independent of semantic understanding, but whereas the meaning of lexical vocabulary varies from context to context (run on the treadmill, run on the market, run-on sentence, etc.) vocabulary has fixed minimal semantic content independent of context.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.005 |
| Scholarly communication | 0.006 | 0.004 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.029 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".