Shelf life: Identifying the abandonment of online digital humanities projects
Bibliographic record
Abstract
Abstract A large portion of the research carried out in the digital humanities has an online digital object (usually referred as a project) as one of its components. In turn, these online digital objects can be catalogued as distributed resources, which implies that the administrative control of information related to a topic may be spread across online resources and/or collections maintained by multiple scholars in different institutions. This administrative decentralization can lead to changes in content that are often unexpected by a researcher, which can be caused by different factors or circumstances. This reasoning led us to formulate the following question: When can online digital humanities projects be considered abandoned? In this article, we carry out a study on the persistence and average life span of online projects in the digital humanities. More specifically, we will elaborate on their reliance on distributed resources and methods for measuring their shelf life: the average length of time that a digital project can endure without updates until it can ultimately be considered abandoned by its researcher.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.015 | 0.007 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".