The Academic Reward System is the Primary Influence Toward Faculty Non-Participation in Institutional Repositories
Bibliographic record
Abstract
Objective – To better understand the lack of faculty participation in Cornell University’s DSpace institutional repository (IR), and to learn if this lack of participation is peculiar to Cornell or reflective of a larger trend in faculty non-participation in IRs. 
 
 Design – Comparative analysis and interviews. 
 
 Setting – Cornell University’s DSpace IR and sciences, social sciences, and humanities faculties; and DSpace installations at 7 other universities.
 
 Subjects – The DSpace IR at Cornell University and at 7 other locations. Eleven sciences, social sciences, and humanities faculty members at Cornell University. 
 
 Methods – The authors analyzed data over a fifteen-month period from Cornell’s DSpace IR to determine the total deposits, the types of objects deposited, the communities and collections that received deposits, the frequency of deposits, the IP addresses which made deposits, and how often objects in the IR were viewed. These data were compared to equivalent data taken from seven other IRs on all aspects except deposits from IP addresses and how often objects were viewed. Finally, 11 Cornell faculty members from various departments in the sciences, social sciences, and humanities were interviewed over a two-month period to provide context to the comparative analysis.
 
 Main results – At the time of the study, the IR at Cornell was organized into 193 communities of collections. These collections numbered 196, with 139 of them holding a combined total of 2646 objects: The other 57 collections were empty. While the IR as a whole showed steady growth, 77% of Cornell’s collections reflected a plateau growth pattern of primarily “one-time deposits,” approximately 18% exhibited a stair-step growth pattern of “periodic batch additions of material,” approximately 3% showed steady growth, and 1.4% were “uncatagorizable.” Five-hundred nineteen unique IP addresses made deposits to Cornell’s IR over the course of the fifteen-month study, but 50% of these deposited only one object, and only 32 IP addresses deposited 10 or more objects. 
 
 Of the other IRs studied, the lowest number of communities is zero and the highest is 390, the number of collections ranged from 10 to 282, and the number of objects ranged from 500 to 32,676. In most statistical categories, Cornell fell in the midrange. The two repositories with the fewest communities and collections – zero communities and 18 collections in one instance, and 6 and 10 in the other – are the only two with no empty collections. The repository with the most communities and collections also had the most empty collections (58%). The repository with the most objects was the one with zero communities and only 18 collections; and the repository with the fewest objects was the one with only 6 communities and 10 collections. The third largest IR, with 3111 objects, had far and away the highest rate of steady growth (16.7%); while the IR with the most objects had the highest rate of stair-step growth (56.3%), and was the only IR to have a higher percentage of growth in any category other than plateau. 
 
 Interviews with faculty indicated that they do not make deposits to IRs for a number of reasons. Faculty considered their primary audience to be their peers, so access to their scholarship was largely considered a “non-issue” as it was adequately provided through personal Web pages, subject repositories, or journal literature. Likewise, long-term preservation was not an overarching area of concern. The chief factors for not using an IR, however, all revolved around restrictions brought on by the academic reward system. Questions of copyright and whether depositing objects qualifies as publishing, thereby hindering efforts to publish in journals, were paramount, as were fears that depositing scholarship alongside less rigorous works in a catch-all IR would diminish the work and the reputation of the scholar by association. Hesitancy to make work available before it had been certified and peer-reviewed was also a foremost concern. 
 
 Conclusion – Although objects in Cornell’s DSpace are accessed both locally for items that are tied into the curriculum, and outside of the university for items that are of national (and international) interest, the repository was not supported well by the faculty. The majority of the collections defined in Cornell’s IR were under populated, and what growth was evident arose primarily from deposits made by non-faculty. The reasons for this were manifold, but centered primarily on the established culture of the academic reward system, which encourages publishing in recognized journals and does little to foster thoughts for long-term preservation or dissemination outside of a given scholar’s peer group. These issues were evident in faculty concerns that depositing materials in an IR might prevent later publication in a journal; the idea that depositing scholarship in a non-vetted repository would diminish that work by association with less scholarly materials; the feeling in some fields that it would be irresponsible to provide access to any unfinished, non-vetted work; the thought that IRs are not sufficient to the task of certifying scholarship; and the concern that deposit in an IR might lead to plagiarism or the loss of initiative on unpublished ideas.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.012 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.120 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".