Bibliographic record
Abstract
Before introducing the 11 interesting and eclectic articles in this general issue, I wanted to talk about IRRODL indexing and improvements in our presentation format and function.Soon after we launched IRRODL in 2001, we were informed by potential authors of the necessity for articles to be published in journals that are included in Thomson's Social Science Citation Index (SSCI), known then as Thomson ISI.In many countries, inclusion in SSCI is mandatory for publications that are valued by institutional and national assessors of research.Since our authors freely license work for publication in this journal, the only reward for the considerable effort involved (besides, of course, untold amounts of fame!) is that the publication should be counted for tenure and promotion and for funding from granting agencies.SSCI is run as a fee service by the Thomson-Reuters publication empire ($13.1 billion annual sales, 55,000 employees) and subscribed to by most academic research libraries.Journals are evaluated for inclusion in SSCI by a formal investigation.One of the rules for inclusion in the index is that the journal must have been publishing regularly for a minimum of three years.Thus, in 2004, I dutifully completed the application form to have IRRODL indexed by SSCI.For years I received no responses to my follow-up emails requesting results of the evaluation.No rejection or failure of the review -just nothing!In 2009, I was able to begin a series of email exchanges with an editor from Thomson-Reuters and this spring we received notice that IRRODL would be indexed beginning with the first issue of 2010!Thus ends a long struggle and one of my favorite rant topics.I still contend that as academic researchers we have given far too much control over our affairs to a commercial publisher.However, I am pleased to be indexed and to have this measure of quality applied to our
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.023 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.006 | 0.005 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.018 | 0.012 |
| Open science | 0.003 | 0.006 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.309 | 0.294 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".