Industry Bibliographical Databases: Perspectives of Use in the FMBA of Russia for Scientific Expertise in Decision-Making. Report 1. General Issues and Database on Health and Other Effects in Nuclear Workers
Bibliographic record
Abstract
The presented review of three reports is devoted to bibliographic databases on health and other effects and indexes in nuclear workers (NW) and uranium miners (U miners), developed within the framework of the research theme of the Federal Medical and Biological Agency of Russia (FMBA) and registered with the state in Rospatent. Report 1 sets out introductory issues of the theory of databases, as well as registers, and provides detailed information on the database for NW. The purpose of the database for NW creating was to form a repository for accessible for abstract and full-text search published data on themes relevant for conducting research examinations for expertise in the system of the FMBA, in other healthcare institutions dealing with the radiation factor, and, more broadly, for conducting fundamental and applied research in the field of professional exposures. The main parts of the database are two separate sub-databases for Russian and foreign NW (Russian NW and Foreign NW), in which the sources are collected in alphabetical order by the authors of the publication or the organizations that created the document. The structural form of information is a catalog that includes primary (main) units of information in the form of an information file about the source (DOC), which contains the title of the publication/document, an abstract (sometimes additional information), and the full original publication (PDF, rarely HTML), available for 88–91 % of sources (the Russian and Foreign sub-databases contain 2078 and 2145 sources, respectively, as of the end of January 2025). 51 % of the works in the database correspond to studies for Russian NW; followed by the USA, Great Britain, Canada, France, and Japan. Visual and/or software search of the material in the database it is supposed to be carried out both through the information names of the catalogs, including the themes of research, carried out using the list of abbreviations (metadata for the database), and through all the texts of the sources included in the database using the proposed programs. Auxiliary elements of the database are fragments of two sub-bases that have undergone hierarchical thematic cataloging in accordance with the identified areas of research on the effects and indexes for NW. These elements are intended, firstly, for initial familiarization with the subject of the database for NW, and, secondly, they are significant as a final thematic base with a certain number of sources, which can be used directly for operational purposes. The developed database for NW has no analogues among industry databases/registers for NW in various countries, nor among bibliographic and search systems. PubMed, Cochrane Library, EMBASE, CINAHL, ISRCTN, Web of Science and Google revealed 5–24 times fewer sources on the theme than the proposed database, and in most cases the world search systems do not provide for the extraction of original publications (as for the IAEA INIS bibliographic database on radiation effects). The depth of the search for works on effects and indexes for NW in the world systems is significantly inferior to the developed database (1960–1970s versus 1940–1950s). It is concluded that the presented database on NW is unique for examination within the framework of the FMBA and other healthcare institutions, and has no complete replacement as a scientific reference and expert depot of sources.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".