Critical Remarks on Reference-Scaled Average Bioequivalence
Bibliographic record
Abstract
PURPOSE: More than a decade ago the option to assess highly variable drugs / drug products by reference-scaled average bioequivalence was introduced in regulatory practice. Recommended approaches differ between jurisdictions and may lead to different conclusions even for the same data set. According to our knowledge, implemented methods have not been directly compared for their operating characteristics (Type I Error and power). METHODS: We performed Monte Carlo simulations to assess the consumer risk and the clinically relevant difference for the recommended regulatory settings. RESULTS: In all methods for reference-scaled average bioequivalence the Type I Error can be inflated with a consequently compromised consumer risk. Furthermore, the clinically relevant difference could vary between studies performed with the same reference product. CONCLUSIONS: Only average bioequivalence with fixed - widened - limits would both maintain the consumer risk and offer an unambiguously defined clinically not relevant difference. As long as such an approach is not implemented in regulatory practice, we recommend adjusting the level of the test a.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.262 | 0.606 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.005 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.002 | 0.017 |
| Scholarly communication | 0.006 | 0.008 |
| Open science | 0.012 | 0.004 |
| Research integrity | 0.019 | 0.034 |
| Insufficient payload (model declined to judge) | 0.007 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".