Bibliographic record
Abstract
“Science” is only one type of knowledge. There are nine faculties at the University of Ottawa; only one is named “science”, and this is typical at a large university. I strongly recommend replacing “science”, “scientists” and “open science” with more inclusive terminology such as “open scholarship” or “open knowledge”, “scholar” or “researcher” in the title and throughout the document. The Pubfair framework is an excellent beginning for a needed profound transformation in how scholars work together and disseminate research. This is the kind of approach most likely to achieve significant savings based on current spend on scholarly publishing, and these savings will be needed to support innovation in scholarly production and dissemination. My recommendation is to proceed with an iterative approach and an initial focus on helping scholarly communities with unmet needs for new forms of review and publishing, such as scholars who create and share datasets or tools using artificial intelligence, digital humanists, and scholarly bloggers. The specific needs for community input whether through review or collaboration in the planning process will vary by discipline and type of product. The work of defining needs and identifying potential solutions should be led by the scholarly community in consultation with repository managers. This is a reversal of the proposed leadership / consultation approach in the framework document. Finally, while I recommend an immediate start to this approach, my advice is to see this as a long-term radical transformation that will likely take decades to complete.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.230 | 0.499 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.004 | 0.003 |
| Bibliometrics | 0.014 | 0.011 |
| Science and technology studies | 0.015 | 0.014 |
| Scholarly communication | 0.044 | 0.013 |
| Open science | 0.011 | 0.010 |
| Research integrity | 0.014 | 0.012 |
| Insufficient payload (model declined to judge) | 0.098 | 0.113 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".