Determining recommended acceptable intake limits for N-nitrosamine impurities in pharmaceuticals: Development and application of the Carcinogenic Potency Categorization Approach (CPCA)
Bibliographic record
Abstract
N-Nitrosamine impurities, including nitrosamine drug substance-related impurities (NDSRIs), have challenged pharmaceutical industry and regulators alike and affected the global drug supply over the past 5 years. Nitrosamines are a class of known carcinogens, but NDSRIs have posed additional challenges as many lack empirical data to establish acceptable intake (AI) limits. Read-across analysis from surrogates has been used to identify AI limits in some cases; however, this approach is limited by the availability of robustly-tested surrogates matching the structural features of NDSRIs, which usually contain a diverse array of functional groups. Furthermore, the absence of a surrogate has resulted in conservative AI limits in some cases, posing practical challenges for impurity control. Therefore, a new framework for determining recommended AI limits was urgently needed. Here, the Carcinogenic Potency Categorization Approach (CPCA) and its supporting scientific rationale are presented. The CPCA is a rapidly-applied structure-activity relationship-based method that assigns a nitrosamine to 1 of 5 categories, each with a corresponding AI limit, reflecting predicted carcinogenic potency. The CPCA considers the number and distribution of α-hydrogens at the N-nitroso center and other activating and deactivating structural features of a nitrosamine that affect the α-hydroxylation metabolic activation pathway of carcinogenesis. The CPCA has been adopted internationally by several drug regulatory authorities as a simplified approach and a starting point to determine recommended AI limits for nitrosamines without the need for compound-specific empirical data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.020 | 0.026 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.006 | 0.003 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".