Predicting patent challenges for small-molecule drugs: A cross-sectional study
Bibliographic record
Abstract
BACKGROUND: The high cost of prescription drugs in the United States is maintained by brand-name manufacturers' competition-free period made possible in part through patent protection, which generic competitors must challenge to enter the market early. Understanding the predictors of these challenges can inform policy development to encourage timely generic competition. Identifying categories of drugs systematically overlooked by challengers, such as those with low market size, highlights gaps where unchecked patent quality and high prices persist, and can help design policy interventions to help promote timely patient access to generic drugs including enhanced patent scrutiny or incentives for challenges. Our objective was to characterize and assess the extent to which market size and other drug characteristics can predict patent challenges for brand-name drugs. METHODS AND FINDINGS: This cross-sectional study included new patented small-molecule drugs approved by the FDA from 2007 to 2018. Market size, patent, and patent challenge data came from IQVIA MIDAS pharmaceutical quarterly sales data, the FDA's Orange Book database, and the FDA's Paragraph IV list. Predictive models were constructed using random forest and elastic net classification. The primary outcome was the occurrence of a patent challenge within the first year of eligibility. Of the 210 new small-molecule drugs included in the sample, 55% experienced initiation of patent challenge within the first year of eligibility. Market value was the most important predictor variable, with larger markets being more likely to be associated with patent challenges. Drugs in the anti-infective therapeutic class or those with fast-track approval were less likely to be challenged. The limitations of this work arise from the exclusion of variables that were not readily available publicly, will be the target of future research, or were deemed beyond the scope of this project. CONCLUSIONS: Generic competition does not occur with the same timeliness across all drug markets, which can leave granted patents of questionable merit in place and sustain high brand-name drug prices. Predictive models may help direct limited resources for post-grant patent validity review and adjust policy when generic competition is lacking.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".