Exploring some intersections between pharmacokinetics, factor <scp>VIII</scp> measurement and human morphometrics – impact of recent advances in haemophilia study design on our understanding of optimal haemophilia treatment
Bibliographic record
Abstract
The interesting paper from Garmann et al. 1 in this issue of Haemophilia provides an opportunity to reflect on progress in understanding how to best define optimal treatment for haemophilia patients. Garmann et al. report their learnings from modelling the pharmacokinetics (PK) of a new factor VIII product (Kovaltry, Bayer Healthcare, Whippany, New Jersey, US) in a sizeable population of severe haemophilia A patients. In the past, the regulatory requirement for approval of new factor concentrates was to assess their bioequivalence with an already approved plasma-derived or recombinant product. A classical bioequivalence study required a minimum of 12 patients, ideally using a cross-over protocol with administration of 50 IU/kg of either product and dense blood sampling to measure factor VIII over at least 32 h, with an optional sample at 48 h 2, 3. More recently, a regulatory requirement mandates preregistration observation for a minimum of 50 exposure days (i.e. approximately 4–6 months) in a larger population of patients (usually 100) to assess safety and efficacy 4-6. Since many of these patients (may) undergo factor VIII measurements during such an observation, a new opportunity for a broader assessment of PK is created, which can be maximized by adopting a population PK approach 7, 8. This statistical approach incorporates dense and/or sparse PK samples from a relatively large population and correlates patient intrinsic (e.g. age, weight, height, race) and extrinsic (e.g. co-morbidities) factors with drug disposition for the enhancement of dosing 8. Depending on the study design, these sparse samples may be measured in a central laboratory, as was the case in the LEOPOLD studies 9, 10, thereby reducing variability in measurements, and strengthening conclusions. Very appropriately, Garmann et al. have not missed this opportunity, and modelled LEOPOLD data with a population PK approach. The lesson learned is the (apparently not so) obvious concept that ‘no measurement done’ is not the same as ‘no measurable factor’ – to paraphrase the famous evidence-based medicine statement about empty systematic reviews: ‘no evidence of effect is not the same as evidence of no effect’ 11. The assay for factor VIII measurement has a lower limit of quantitation of 0.01 IU·mL−1 (or 1 IU·dL−1) for many clinical laboratories. Values below this level are reported as ‘<0.01 IU·mL−1’ which, in PK jargon, constitutes a ‘below limit of quantitation’ (BLQ) measurement 12. The level of 0.01 IU·mL−1 is of clinical importance in that it defines patients with severe haemophilia (those with baseline factor level below 0.01 IU·mL−1), and it is a founding concept of prophylaxis. Persons living with severe haemophilia experience multiple spontaneous bleeds, whereas those with higher baseline levels (i.e. =>0.01 IU·mL−1) do not. The goal of prophylaxis has historically been defined as aiming to keep plasma factor levels above the critical concentration of 0.01 IU·mL−1 13, a goal that can be verified by drawing blood samples for factor VIII measurements before the following infusion (trough levels). On clinical grounds, trough levels provide very important information regardless of whether factor VIII is measurable or not (i.e. BLQ). Imagine two patients, one of whom has 0.025 IU·mL−1 at 72 h, and the other <0.01 IU·mL−1. You have no doubts that they are different, no matter if both had a factor VIII of 0.05 IU·mL−1 48 h following the infusion. These patients require different treatment approaches (the former may go well with every third day, the latter may require every other day infusions), and these approaches can be defined by drawing samples at times when BLQs may occur. The novelty of Garmann et al. paper is to show that BLQs provide precious information also for PK modelling. Many of the sparse samples in the LEOPOLD studies were trough levels and BLQs, and the authors clearly showed that incorporating the information provided by those BLQ samples is critical to a population PK approach for factor VIII. Not considering the proportion of samples (and patients) at BLQ leads to a gross underestimate of clearance, overestimate of the terminal half-life and overestimate of the time spent above critical concentrations. One may wonder if the same applies to classical PK estimates of factor concentrates. Yes, it does, and with little to no chance of detecting and correcting the problem when estimating PK without late samples. Has consideration of BLQs always been done in published PK estimates for factor VIII? This is difficult if not impossible to tell from many reports. Is it easy to do? It is doable but incorporating ‘<0.01 IU·mL−1’ in regression models is not as straightforward as one could imagine 14. When comparing different factor VIII products based on published PK estimates (obtained by either classical or population PK), one might want to check whether BLQs have been used or not; if not, then half-lives may have been overestimated, often resulting in a more favourable profile for the factor assessed disregarding BLQs. Similarly, when assessing the PK of individual patients, one wants to use all data, including measurements resulting in BLQs. No matter if using a classic or population PK approach to individual PK estimation, one wants to be sure BLQs are appropriately used. Secondly, total body weight is not the best way to dosing factor VIII. What Garmann et al. have found is that the most appropriate covariate in modelling data from the LEOPOLD study is lean body weight (LBW). This is not the first time that a measure of body mass composition has been found to be better than total body weight in explaining inter-individual variability in factor VIII PK 15-18. LBW, adjusted body weight or body mass index (BMI) have all been correlated with PK of factor VIII. The consequence of this is that one needs to carefully choose the body mass measure to adopt in modelling PK and ensure that appropriate morphometrics are measured in the study protocol (e.g. total body weight and height). More importantly, one needs to consider switching from clinical regimens built on total body weight to more efficient approaches. In oncology, e.g. many drugs are dosed on body surface area (BSA), and it would be possible switching to dose based on LBW in haemophilia. This might indeed reduce the variability observed in clinical practice, and help get more patients in the proper therapeutic range (avoiding undesirable high peaks and low troughs). Finally, due consideration is to be given to the uncertainty involved in designing treatment regimens for haemophilia patients. A first layer of uncertainty lies in the laboratory measurements themselves. When considering a laboratory result, e.g. a 0.4 UI mL−1 of factor VIII, one tends to forget that measurement error is real. Measurement error for clotting assays is usually not <10%, meaning that the 0.4 can be anywhere from 0.36 to 0.44 (in terms of the old ‘percent’, result can be anywhere between 36 and 44%). Measurement error is incorporated into the residual variability in the Garmann et al. model, which is then used during post hoc individual PK estimation 7. A second layer is the variability related to the specific factor assay 19; using different partial thromboplastin time (PTT) reagents, 0.4 IU·mL−1 could have been, let us say, 0.21 or 0.55 UI·mL−1. This second layer is a particularly relevant problem when measuring the new extended half-life products, which show a significantly different response to different PTT reagents. A third layer is the variability in the PK estimation process for individual patients. There is a general understanding in PK that there is no perfect model, method or approach – PK modelling is the result of a series of educated guesses about the best representation of an underlying unknown process. The totality of human variability is impossible to capture in one model and therefore, the use of the resulting model may not be appropriate for every patient. What can be expected is that as one gathers more information and develops new methods and techniques, model performance will increase. This is likened to competitive sports, where better training and materials have led to continuous improvement in world records. Garmann et al. have raised the bar for the way we do PK in haemophilia. AI's Institution has received project based funding via research or service agreements with Bayer, Biogen Idec, Grifols, NovoNordisk, Octapharma, Pfizer and Shire (formerly Baxter and Baxalta). ANE has no disclosures
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.151 | 0.148 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.001 | 0.007 |
| Scholarly communication | 0.005 | 0.008 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.002 | 0.006 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".