Development of a preference-based measure for Multiple Sclerosis: the Preference-Based Multiple Sclerosis Index (PBMSI)
Bibliographic record
Abstract
Assessing health-related quality of life (HRQL) has moved to the forefront of clinical research andis considered a crucial endpoint of clinical interventions. One approach to assessing HRQL isthrough the use of health profiles. Health profiles are analyzed by sub-scale, where each sub-scalerepresents a domain of health. These measures do not provide information on the relativeimportance attached to each domain. As a result, the domains cannot be combined into an overallscore, and a trade-off cannot be made between domains when evaluating the effectiveness ofinterventions. Another approach to measuring HRQL is through the use of preference-basedmeasures. Not only do these measures provide descriptive information on the various dimensionsof health, but also provide a value for each. They have the advantage of leading to a single numberthat balances gains in one domain against losses in another. When linked to life-expectancy, theyprovide measures of quality adjusted life years (QALY) and are used to make decisions about thecost-effectiveness of interventions. The best known preference-based measures are the HealthUtilities Index (HUI), the EuroQol-5D (EQ-5D) and the Short Form-6D (SF-6D). However, thechallenge of using such generic preference-based measures in people with Multiple Sclerosis (MS)is that they may not capture all domains of health relevant to the disease and the domain weightingis based on the values from the naive general population.Therefore, the overall objective of this PhD thesis is to take important steps towards developing aPreference-Based Multiple Sclerosis Index (PBMSI) for use as a global outcome in clinical andcost-effectiveness studies for MS.To do this, a systematic review of HRQL outcomes in MS interventions was carried out andidentified that an imporant source of heterogeneity in the literature arises from the many differentmeasures used and domains evaluated (Manuscript 1). As preference-based measures reduce someof the heterogeneity by yielding one value across mutliple domains of health, the content ofgeneric preference-based measures was assessed in light of the domains identified as beingimportant to people with MS (Manuscript 2), and a review of their psychometric properties wascarried out (Manucript 3). Results revealed that these generic measures were missing severaldomains that were affected by MS, such as walking, fatigue and cognition, identifying ameasurement gap. Making use of a rich data source (that I had previously collected as part of myMSc), optimally performing items targeting the important MS domains were identified and tested10for their discriminatory capacity with respect to known groups with differing disability(Manuscript 4). This study yielded a set of 5 bilingual items (English and French) ready for testingfor comprehension and wording using cognitive interviewing with a sample of 22 people with MS(Manuscript 5). An item met criteria for acceptability after 3 to 4 rounds of interviews.The final step in this thesis was to elict preferences for different health states generated throughcombinations of items, using two different standard methods of preference elicitation which areknown to have conceptual and practical differences (Standard Gamble and Rating Scale).Manuscript 6 presents the results of this preliminary investigation in a sample of 61 patients withMS. The results indicate that the Standard Gamble is difficult for patients to understand andproduces higher values than the Rating Scale. The scoring algorithm developed based on each ofthe methods yielded vastly different results. Although the Standard Gamble is a classical techniqueof measuring preferences using decision making, it was not practical in this patient population. Onthe other hand, the Rating Scale is more suitable for the population but the values are not choicebased potentially limiting their use for economic evaluation of interventions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.037 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".