A Hierarchical Approach to Coding Chemical, Biological and Pharmaceutical Substances
Bibliographic record
Abstract
This hierarchical coding system is designed to classify substances into successively subordinate categories on the basis of chemical, physical and biological properties. Although initially developed for occupational cancer epidemiological studies, it is general in nature and can be used for other purposes where a systematic approach is needed to catalogue or analyze large numbers of substances and/or physical properties. The coding system incorporates a multi level approach, where substances can be coded both on the basis of function and composition. On the first level, a three digit code is assigned to each substance to indicate its primary use in the occupational environment (e.g. pesticide, catalyst, adhesive). Substances can then be coded using a ten digit code to indicate structure and composition (e.g. organic molecule, biomolecule, pharmaceutical). Depending on the complexity required, analysis can incorporate the three digit code, ten digit code, or a combination of both. The approach to coding both chemical and biological agents is modeled in part after conventional approaches used by the International Union of Pure and Applied Chemists (IUPAC) and the International Union of Biochemists (IUB). Development of the coding system was initiated in the 1980's in response to a need for a system allowing analysis of individual agents as well classes or groups of substances. The project was undertaken as a collaborative venture between the BC Cancer Agency, Cancer Control Research program (then Division of Epidemiology) and the Department of Chemical and Biological Engineering at the University of British Columbia.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".