World War I and Siberia: on the statistics, typology and range of topics of print publications
Bibliographic record
Abstract
The article considers the character of print publications related to the topic “World War I and Siberia”. The research deals with works published in the Siberian region in 1914-1917 and those published in the first decade of the XXI century. Different statistical, typological and thematic characteristics of printed works are presented, which is reflected in a number of tables. 123 works on the subject “World War I and Siberia” were published in 12 cities of the region in 1914-1917. A significant percentage of such works was published in Omsk (nearly a quarter of all publications). The publications represented various types of literature including materials related to production, education, propaganda and popular science, as well as reference and fiction books. In the first decade of the XXI century, 238 works on the topic were published in 31 cities of Siberia, the Russian Far East, the Ural region and the European part of Russia. The biggest number of them was still published in Omsk. Almost all works on the topic were of scientific nature. The largest number of papers were published as scientific conference proceedings. The works covered a variety of issues. Particular attention was paid to the development of the regional economy, the issues of prisoners of war and detainees, as well as to the functioning of local self-governance bodies. Based on the comparison of works published during the World War I and the first decade of the XXI century, the researcher points out that typologically the two groups were of different nature and the thematic areas also differed. In addition, the literature published in 1914-1917 can serve as a source base for solving problems of a historiographical nature.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.035 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.052 | 0.105 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.013 | 0.006 |
| Open science | 0.001 | 0.006 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.007 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".