Bibliographic record
Abstract
The Smithsonian/NASA Astrophysics Data System (ADS) is the core digital library for astrophysics. It was built in 1992, put on the Internet for public use in 1993, and moved to the WWW in 1994. It was an immediate success, and has dominated astronomers use of their scholarly literature ever since. By the beginning of the 21<sup>st</sup> century it was clear the ADS was an important part of the scholarly infrastructure of astronomy. This required that it be rebuilt, from the original <em>ad hoc</em> system to more long-term stable, maintainable, robust and extensible one, while keeping the existing service operating 24/7/365. This has been substantially more difficult than we imagined; finally begun in 2007 the new system was released this spring (2018), to run for a year in parallel with the old, Classic system until next spring, when the Classic system will be turned off, after a quarter of a century of use. The project has taken more than a decade of intense work; it required a change in the management structure, and an eventual doubling of the staff (and thus budget). During this time the technological landscape has changed substantially, many promising directions for the development were taken, only to be later abandoned. Additionally many new capabilities became possible, requiring improvements both to the new design, and where feasible, to the old system as well. We hope the new ADS will enable researchers discovery for the next 25 years. The ADS is at ads.harvard.edu
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.338 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".