Trends in covalent drug discovery: a 2020–23 patent landscape analysis focused on select covalent reacting groups (CRGs) found in FDA-approved drugs
Bibliographic record
Abstract
INTRODUCTION: Covalent drugs contain electrophilic groups that can react with nucleophilic amino acids located in the active sites of proteins, particularly enzymes. Recently, there has been considerable interest in using covalent drugs to target non-catalytic amino acids in proteins to modulate difficult targets (i.e. targeted covalent inhibitors). Covalent compounds contain a wide variety of covalent reacting groups (CRGs), but only a few of these CRGs are present in FDA-approved covalent drugs. AREAS COVERED: This review summarizes a 2020-23 patent landscape analysis that examined trends in the field of covalent drug discovery around targets and organizations. The analysis focused on patent applications that were submitted to the World International Patent Organization and selected using a combination of keywords and structural searches based on CRGs present in FDA-approved drugs. EXPERT OPINION: A total of 707 patent applications from >300 organizations were identified, disclosing compounds that acted at 71 targets. Patent application counts for five targets accounted for ~63% of the total counts (i.e. BTK, EGFR, FGFR, KRAS, and SARS-CoV-2 Mpro). The organization with the largest number of patent counts was an academic institution (Dana-Farber Cancer Institute). For one target, KRAS G12C, the discovery of new drugs was highly competitive (>100 organizations, 186 patent applications).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".