Proceedings of the 2017 Web Archiving and Digital Libraries Workshop
Bibliographic record
Abstract
1. HTTPreserve – Web Preservation in Documentary Heritage by Ross Spencer 2. WARC-Portal: A Tool for Exploring the Past by Muhammad Umar Qasim 3. Impact of URI Canonicalization on Memento Count by Mat Kelly, Lulwah M. Alkwai, Sawood Alam, Michael L. Nelson, and Michele C. Weigle and Herbert Van de Sompel 4. Web Archiving Through In-Memory Page Cache by Saket Vishwasrao, Zhiwu Xie, and Edward A. Fox 5. Building a National Web Archiving Collaborative Platform: The Web Archives for Longitudinal Knowledge Project by Ian Milligan, Nick Ruest, and Ryan Deschamps 6. Avoiding Zombies in Archival Replay Using ServiceWorker by Sawood Alam, Mat Kelly, Michele C. Weigle, and Michael L. Nelson 7. Topic Shifts Between Two US Presidential Administrations by Ziquan Wang, Borui Lin, Ian Milligan, Jimmy Lin 8. Web archives: A preliminary exploration of user expectations vs. reality by Brenda Reyes Ayala 9. Challenges for Grassroots Web Archiving of Environmental Data by Emily Maemura, Dawn Walker, Matt Price, and Maya Anjur-Dietrich 10. Legal Deposit, Collection Development, Preservation, and Web Archiving at Library and Archives Canada by Tom Smyth 11. Working Together Toward a Shared Vision: Canadian Government Information Digital Preservation Network (CGI DPN) by Muhammad Umar Qasim and Sam-Chin Li 12. Strategies for Collecting, Processing, and Analyzing Tweets from Large Newsworthy Events by Nick Ruest 13. Classification of Tweets using Augmented Training by Saurabh Chakravarty, Eric Williamson, and Edward Fox
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".