Bibliographic record
Abstract
Introduction Language is an immensely rich phenomenon, presenting vast challenges for the linguist, the scientist of this phenomenon. Among the most central, and most difficult, questions are ones concerning the nature of the capacity we all have to speak a language. Just what is this capacity, and how does it arise in the individual? We thus have (among many others) the two following related questions: - What is the correct characterization of someone who “knows a language” (in general terms, who has command of a systematic connection between sound and meaning)? - How does that systematic connection arise in the individual? For the first of these questions, the linguist hopes to account, in an explicit way, for the speaker’s ability to put together and understand sentences, including ones new to the speaker (and often new to all speakers), and for the speaker’s ability to judge potential sentences as “acceptable” ( John left ) or “unacceptable” ( Left John ). For the second question, a particularly difficult one given how complex the capacity seems to be and how quickly it is acquired, the linguist seeks to discover what aspects of the capacity are determined by the child’s experience, and how this determination (“language learning”) takes place. For a half century, Noam Chomsky has been developing a theory of language that deals with these two questions, by positing explicit formulations of human language capacity in terms of a productive “computational” system, most of whose properties are present in advance of experience, “wired in” in the structure of the human brain. Thus, Chomsky conceives of his enterprise as part of psychology, ultimately biology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.021 |
| Scholarly communication | 0.007 | 0.006 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.008 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".