Bibliographic record
Abstract
The World Wide Web (WWW) was formally introduced as a proposal in March 1989 (http://info.cern.ch/Proposal.html) and implemented in May 1990 (http://info.cern.ch/) by Sir Tim Berners-Lee of CERN (European Organisation for Nuclear Research). The novelty of the concept proposed was in its hypothetical capacity to share information easily over the Internet by deploying hyperlinked hypertext, encoded, displayable and retrievable through hypertext transfer protocol (HTTP) and hypertext markup language (HTML), by means of a browser. The Internet, a global system of interconnected computer networks based on the TCP/IP communications protocol standard, predated the Web by approximately thirty years. The “Web”, which would eventually become the most widely used portion of this broader Internet, was made available to the public in 1991. The first Web browser with a graphical user interface (GUI), Mosaic, was introduced in 1993 and from this point on enabled users to interface more intuitively with the Web via icons and visuals rather than text commands. Technically, in a period of just 20 years, the Web has evolved from an information repository of posted static text Web pages to a dynamically charged user-interactive environment (“Web 2.0”) of social networking sites, multimedia content-sharing sites, and real-time communication propelled on the back-end by diverse types of specialized servers, and database and content management systems (CMS). Web technologies and standards, linked to industry initiatives, academic research projects, and international organizations such as the W3C, Unicode Consortium and ICANN, continue to evolve rapidly.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.017 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.007 |
| Science and technology studies | 0.004 | 0.007 |
| Scholarly communication | 0.016 | 0.015 |
| Open science | 0.002 | 0.007 |
| Research integrity | 0.004 | 0.004 |
| Insufficient payload (model declined to judge) | 0.178 | 0.128 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".