Partially hidden multi-state modelling of a prolonged disease state defined by a composite outcome
Bibliographic record
Abstract
For rheumatic diseases, Minimal Disease Activity (MDA) is usually defined as a composite outcome which is a function of several individual outcomes describing symptoms or quality of life. There is ever increasing interest in MDA but relatively little has been done to characterise the pattern of MDA over time. Motivated by the aim of improving the modelling of MDA in psoriatic arthritis, the use of a two-state model to estimate characteristics of the MDA process is illustrated when there is particular interest in prolonged periods of MDA. Because not all outcomes necessary to define MDA are measured at all clinic visits, a partially hidden multi-state model with latent states is used. The defining outcomes are modelled as conditionally independent given these latent states, enabling information from all visits, even those with missing data on some variables, to be used. Data from the Toronto Psoriatic Arthritis Clinic are analysed to demonstrate improvements in accuracy and precision from the inclusion of data from visits with incomplete information on MDA. An additional benefit of this model is that it can be extended to incorporate explanatory variables, which allows process characteristics to be compared between groups. In the example, the effect of explanatory variables, modelled through the use of relative risks, is also summarised in a potentially more clinically meaningful manner by comparing times in states, and probabilities of visiting states, between patient groups.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".