AMCIS 2008 Panel Report: Aging Content on the Web: Issues, Implications, and Potential Research Opportunities
Bibliographic record
Abstract
Since its inception in the early 1990s, the World Wide Web (Web) has grown enormously. According to the “official Google blog” (Google 2008), the Web had 1 trillion (as in 1,000,000,000,000) unique coexisting URL’s as of July 25, 2008. Given the exponential growth of the Web over time, an issue that is likely to gain prominence is that of outdated information. This is especially important to study since many of us rely on the Web to find facts in order to take decisions. For example, for students and researchers, the “date” of a document is important for scholarship and student work. However, getting an accurate date on content is challenging, and furthermore, outdated pages that are not deleted from Web servers will continue to be returned in response to Web searches. The panel, held at the 2008 Americas Conference on Information Systems in Toronto, Canada, identified a number of research issues and opportunities that arise as a result of this phenomenon.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.021 | 0.024 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.005 |
| Science and technology studies | 0.005 | 0.002 |
| Scholarly communication | 0.008 | 0.005 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.027 | 0.010 |
| Insufficient payload (model declined to judge) | 0.016 | 0.008 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".