Bibliographic record
Abstract
Historically, privacy was almost implicit, because it was hard to find and gather information. But in the digital world, whether it's digital cameras or satellites or just what you click on, we need to have more explicit rules – not just for governments but for private companies. Bill Gates (1955–) Wired.com , 12 November 2013 When an archivist receives archival materials into her care her first interest is, almost inevitably, to make the documentary treasures available for use. What a thrill it is to hold in your hands an original letter written by your favourite artist, or to see the name of a former mayor among the rolls of students in your local high school. As someone once said, archivists get paid to read other people's mail. We love the opportunity to see into the lives of other people and, through them, to understand the way our communities functioned in years past. But the artist and the mayor have rights too. And these rights do not vanish when documents created or owned by them move into archival custody. The need to balance access and privacy and the need to respect intellectual property rights are perhaps the greatest source of tension for the archivist, especially given the ubiquity of digital technologies today. The requirements of access, privacy and copyright laws can throw obstacles in the way of achieving the goal of archival service, which, as stated many times already, is to support the acquisition, preservation and management of archival materials so that they can be made available for use. But the creators of documentary materials – the people who kept personal diaries, wrote letters to their sister, took photographs on their holiday, prepared financial reports for their business – did not create those records for posterity. The fact that those records ended up in an archival repository, perhaps decades after their creation, does not mean that the creators of those records have lost all right to control the ways in which those materials may be used.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.004 | 0.016 |
| Scholarly communication | 0.015 | 0.025 |
| Open science | 0.002 | 0.010 |
| Research integrity | 0.005 | 0.007 |
| Insufficient payload (model declined to judge) | 0.021 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".