Analysis of the structural diversity of heterocycles amongst European medicines agency approved pharmaceuticals (2014–2023)
Bibliographic record
Abstract
This review presents a detailed analysis of the heterocycle diversity amongst medicines with new active substances (NAS) approved by the European Medicines Agency (EMA) in the 10 years from 2014-2023. A total of 380 medicines were approved that contain a NAS, of which 160 are small molecule products that contained one or more NAS with a heterocycle (164 NAS in total). Of the 164 heterocycle-containing NAS, 76% contained more than one heterocycle. The majority (59%) of the 164 active substances contained at least one fused heterocycle. The most common bicyclic rings were quinoline, benzimidazole, indole, and pyrrolopyrimidine. Tricyclic and polycyclic fused rings were observed but were rare. There were 28 distinct monocyclic heterocycles, consisting of 3, 4, 5, and 6 membered rings. 5-Membered rings were the most diverse as 15 of the 28 heterocycles are 5-membered rings. 6-Membered rings ranked second with 12 heterocycles. There was one 3-membered ring and one 4-membered ring seen. Nitrogen was by far the most common heteroatom in both monocyclic and fused heterocycles. Oxygen, sulfur and boron appeared in monocyclic heterocycles, whilst oxygen, sulfur and phosphorous were noted in fused heterocycles. The most common monocyclic heterocycles were pyridine, piperidine, pyrrolidine, piperazine, pyrimidine, pyrazole, triazole, imidazole and tetrahydropyran. This analysis provides valuable information on the structural diversity of heterocycles that were present in EMA approved medicines between 2014-2023. It highlights heterocycle occurrences, diversity, substitution patterns, and trends. The information detailed will be of interest to organic chemists, researchers, regulatory agencies, and the pharmaceutical industry as it demonstrates how common heterocycles are seen amongst EMA approved medicines for a wide range of therapeutic areas.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".