Gendered effects of pay for performance among family physicians for chronic disease care: an economic evaluation in a context of universal health coverage
Bibliographic record
Abstract
BACKGROUND: Despite increasing popularity among health organizations of pay for performance (P4P) for the provision of comprehensive care for chronic non-communicable diseases, evidence of its effectiveness in improving health system outcomes is weak. An important void in the evidence base is whether there are gendered differences in P4P uptake and in related outcomes amenable to healthcare improvement. This study assesses the gender-specific effects of P4P among family physicians on diabetes healthcare costs in a context of universal health coverage. METHODS: We use population-based linked longitudinal administrative datasets on chronic disease cases, physician billings, hospital discharge abstracts, and physician and resident registries in the province of New Brunswick, Canada. We estimate the effects of introduction of a P4P scheme on excess public healthcare costs among cohorts of adult diabetes patients using propensity score-adjusted difference-in-differences regressions stratified by physician's gender. RESULTS: We observed greater male physician uptake of incentive payments, seemingly exacerbating gender gaps in professional remuneration. Regression results indicated P4P did not lead to improved outcomes in terms of preventing hospitalization costs among patients, only measurable increases in compensation for both the male and female physician workforce. CONCLUSIONS: While P4P was not attributed in this study to reduced hospital burden and enhanced sustainability of healthcare financing, incentive payments were found to be related to earning gaps by physician's gender. Decision-makers should consider that benefits of P4P be monitored not only for patient metrics but also for provider metrics in terms of gender equality especially given feminization of primary care medical workforces.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".