Bibliographic record
Abstract
In 2007, authors and publishers released book manuscripts with a value of $9.1 billion. These books will continue to be sold for decades. Because of this long working life, the international guidelines for national accounts recommend that countries classify production of books and other entertainment, literary, and artistic originals as an investment activity and then depreciate those books over time. However, the Bureau of Economic Analysis did not capitalize this category of intangible assets until the July 2013 benchmark revision. In order to change the national accounts, this paper collects data on production of book manuscripts from 1900 to 2010. The paper then calculates how GDP statistics change when book manuscripts are classified as capital assets. The main empirical results are as follows: 1) Book manuscripts have a useful lifespan of at least fifty years with an annual depreciation rate of 12 per cent per year; 2) Nominal book investment has hovered around 0.06 per cent of nominal GDP from the 1930s until 2010. Therefore, nominal GDP growth does not change much when book production is classified as an investment activity; 3) After 1970, book investment prices rose faster than overall GDP prices. Accordingly, average inflation rises slightly when book production is classified as investment. Before 1970, book investment prices roughly track overall GDP prices.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.004 | 0.010 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.008 | 0.005 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.046 | 0.017 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".