Legal Protocols and Practices for Managing Copyright in Electronic Theses
Bibliographic record
Abstract
At Queensland University of Technology (QUT) in Brisbane Australia, PhD and Masters by Research candidates are required to deposit both print and digital copies of their theses and dissertations. The fulltext of these digital theses is then made freely available online via the Australian Digital Thesis (ADT) collection. Management of copyright issues has been a major headache and workload problem for the Library: there are many parties involved in the deposit process, and the lack of a common understanding about the rights and responsibilities of the various stakeholders has made the process very complex and time consuming. The response of some universities has been to limit access to just the metadata and abstract. At QUT this is not an option as the University is committed to freeing up access to publicly funded research and its outputs. QUT is also at the forefront of various open access initiatives, including the Open Access to Knowledge Law (OAK Law) Project that is working to develop legal protocols for managing copyright issues in an open access environment and investigate provision and implementation of a rights expression language for implementing such protocols. This paper discusses the aims of this project in relation to the QUT ETD experience, as well as how these fit in with the larger ETD open access environment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.314 | 0.335 |
| Meta-epidemiology (narrow) | 0.001 | 0.003 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.013 | 0.012 |
| Science and technology studies | 0.018 | 0.035 |
| Scholarly communication | 0.046 | 0.041 |
| Open science | 0.010 | 0.021 |
| Research integrity | 0.019 | 0.014 |
| Insufficient payload (model declined to judge) | 0.010 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".