Patterns of similarity and difference between the vocabularies of psychology and other subjects.
Bibliographic record
Abstract
The vocabulary of Anglophone psychology is shared with many other subjects. Previous research using the Oxford English Dictionary has shown that the subjects having the most words in common with psychology are biology, chemistry, computing, electricity, law, linguistics, mathematics, medicine, music, pathology, philosophy, and physics. The present study presents a database of the vocabularies of these 12 subjects that is similar to one previously constructed for psychology, enabling the histories of the vocabularies of these subjects to be compared with each other as well as with psychology. All subjects have a majority of word senses that are metaphorical. However, psychology is not among the most metaphorical of subjects, a distinction belonging to computing, linguistics, and mathematics. Indeed, the history of other subjects shows an increasing tendency to recycle old words and give them new, metaphorical meanings. The history of psychology shows an increasing tendency to invent new words rather than metaphorical senses of existing words. These results were discussed in terms of the degree to which psychology's vocabulary remains unsettled in comparison with other subjects. The possibility was raised that the vocabulary of psychology is in a state similar to that of chemistry prior to Lavoisier.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.015 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.006 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.000 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".