An Approach to Evaluate the Costs and Outputs of Academic Biobanks
Bibliographic record
Abstract
Academic biobanks commonly report sustainability challenges, which may be exacerbated by a lack of information on biobank value. To better understand the costs and supported outputs that contribute to biobank value, we developed a systematic, generalizable methodology to determine biobank inputs and publications arising from biobank-supported research. We then tested this in a small cohort ( n = 12) of academic cancer biobanks in New South Wales, Australia. A proforma was developed to capture monetary and in-kind biobank costing data from biobank managers and publicly available sources. Participating biobanks were grouped and compared according to the following two classifications: open- versus restricted-access and high versus low total annual costs. Our methodology provides a feasible approach for capturing comprehensive costing data for a defined period. Characterization of biobanks using this approach showed that median total costs, as well as median staffing and in-kind costs, were comparable for open- and restricted-access biobanks, as were the quantity and journal impact metrics of supported publications. High- and low-cost biobanks supported similar median numbers of publications; however, high-cost biobanks supported publications with higher median journal impact factor and Altmetric scores. Overall, 9 of 10 biobanks had higher Field-Weighted Citation Impact scores than the global average for similar publications. This is the first tested, generalizable approach to analyze the costs and publications arising from biobank-supported research. By determining explicit cost and output data, academic biobanks, funders, and policymakers can engage in or support informed redirection of resourcing and/or benchmark setting with the aim of improving biobank support of research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.050 | 0.142 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.020 | 0.025 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.005 |
| Open science | 0.002 | 0.005 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".