Bibliographic record
Abstract
70% of the participating institutes have one research repository, nearly a quarter more than one. Approximately 8% have outsourced it. The large majority of research repositories is focused on containing various publication types in full text, while less than half also contains metadata-only records relating to publications. But quantitatively, metadata-only records take up 51% of all records, while full text records take up one-third. Only a small percentage of the repositories contain non-textual materials such as primary datasets, images, video and music. A closer look at the full text records in the repositories shows that 62% are grey literature, i.e. theses, proceedings and working papers; 38% contain primary literature, i.e. journal articles and books/book chapters. The respondents also estimated the number of records of each type in their repositories. From these data it appears that a typical research repository in Europe contained in total 8.545 items in September 2008. With regard to the full text of elsewhere-published materials, one is confronted with copyright rules and, in case of journal articles, the question of which version should be deposited. The comparison of the data from the 2008 and 2006 surveys show a clear trend from preprint form and/or published form towards postprint form. Another important issue is the variation in availability forms for full text supported by the repository: Open Access, Open Access with embargo period, Campus Access or No Access (archive only). It appears that the repositories are offering more options over the last few years. However, from an additional analysis of the 2008 data, it appears that 47% of the research repositories still offer only one form of availability, the Open Access option. The disciplines are fairly even represented in the materials, be it with a slight overrepresentation of Humanities and Social Sciences with 35%. Comprehensive coverage is an important success factor for the research repositories. Coverage is estimated by the respondents of the 2008 survey on average at 35% of the research output of their institutes. In 42 another estimate by the respondents, the percentage of academics of their institutes delivering material to the research repositories is on average 33%. These estimates are similar to those made by the respondents to the 2006 survey, suggesting no real progress in this respect. Work processes vary from self-depositing by academics to independent collection of the materials by repository staff members. Compared to the results of the 2006 survey, there is a remarkable increase in the percentage of repositories that use a combination of various workflows (28% in 2006 versus 44% in 2008).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".