Application of Probabilistic Linkage: Compare Health Care Costs among Menopausal Women with Different Symptoms by Linking Women’s Registry & Claims Data
Bibliographic record
Abstract
Objectives: Menopause symptoms are a good disease severity proxy for menopausal women, but are not available in claims database. We applied probabilistic linkage to add symptoms recorded in a registry database to claims data, and compare the healthcare costs among women with various symptoms. Methods: Women age 45 or older who used estrogen only hormone therapy (HT) were selected from a large U.S. claims database (04/01/2005-09/30/2008). Another group who used estrogen only HT with a menopause diagnosis was selected from the University of Michigan Women’s Registry Database. Logistic regression was used to calculate the propensity score for each patient controlling for osteoporosis, gynecological disorders/procedures, genital infection, gynecology system cancer, breast condition, gut condition, hormone disorder, nerve problem, and other individual comorbidities such as rheumatoid disease, depression, and blood clotting. Patients with the closest propensity score from each group were matched, and menopause symptoms for registry patients were added to the claims database records. After repeating probabilistic linkage 250 times, the mean and 95% confidence interval (CI) of healthcare costs during the follow-up period were calculated. Results: 80 patients from each population were matched after probabilistically linking 20,020 claims database patients with 83 registry database patients. The average cost of patients with at least one symptom was much higher than for patients without symptoms ($13,570 [95% CI: $13,459-$13,680] vs. $3,391 [95%CI: $3,345-$3,436], p-value<0.001). (1 US Dollar= 0.75 Euro) Cost differences were mainly from inpatient, physician visit, and pharmacy costs. Among patients with menopause symptoms, those with hot flashes had the highest costs ($10,127), followed by memory loss ($1,653), vaginal dryness ($864), reduced libido ($568), and mood swings ($358). Conclusions: Women with menopause symptoms incur higher healthcare costs than those without This study suggests symptoms are important determinants of healthcare expenses and their impact can be assessed by linking registry and claims databases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.043 | 0.092 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.004 | 0.005 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".