Pay-for-Performance Reimbursement in Health Care: Chasing Cost Control and Increased Quality through "New and Improved" Payment Incentives
Bibliographic record
Abstract
Abstract Pay-for-performance (P4P) reimbursement has become a popular and growing form of health care payment built on belief that payment incentives strongly affect medical providers' behavior. By paying more to those providers who are deemed to deliver better care, goal is to increase and, hopefully, restrain cost growth. This article provides a brief explanation of: (1) how previous P4P plans in U.S. have fared, along with their special relationship to primary care, and (2) how England's experience with P4P and newer versions of these kinds of plans being pursued in places such as Massachusetts might provide valuable case studies for how U.S. and other countries can achieve meaningful reform of health care organization, delivery and finance. Background, Performance of Early Plans, and Primary Care P4P financially rewards medical providers who achieve, improve upon or exceed performance goals on specified benchmarks. It has developed largely in response to cost control problems and perverse incentives associated with fee-for-service reimbursement, which is dominant model in U.S. (1) Instead of simply reimbursing providers more for greater volume and intensity of care, P4P pays more to providers whose care is deemed to be of higher (or sufficiently high) quality. (2) These plans are intended to lower health care costs over long term by increasing preventative care, primary care and improved treatment of conditions at earlier stages of development. (3) Most P4P approaches adjust payments to hospitals, individual physicians, networks of physicians or medical practice groups in one of three ways: (1) a bonus payment based on a percentage of all care delivered by a provider, (2) a bonus payment per patient member for a provider that has delivered what pre-determined measures would deem as high quality care, or (3) as a percentage of total cost savings achieved relative to what costs would have been without achieving higher quality. (4) The first generation of P4P plans that proliferated from early to mid-2000s in U.S. proved mostly ineffective in either increasing or controlling costs. (5) The bonus payments were arguably too small and areas of clinical too narrow to foster significant behavioral change on part of providers. (6) Moreover, complex patients with multiple medical problems posed unique dilemmas for physicians when their complicated conditions did not fit neatly within individual care guidelines, (7) and their care was (often minimally) coordinated among different clinicians. (8) Concerns emerged that the methods used to measure of care unfairly penalized providers caring for patients with multiple chronic conditions. (9) Studies found that some P4P plans did actually worsen existing disparities by discouraging physicians from caring for poorer, less compliant patients. (10) In short, some providers began cherry picking to avoid those potential patients who they perceived were likely to lower their overall scores. (12) One of most prevalent changes associated with early P4P plans was increased documentation. (13) In other words, rather than increases in and use of preventive services, early P4P plans generated more record-keeping. If pay for performance was a therapy, an observer noted in 2007, its rapid diffusion thus far would have to be considered premature. (14) One of areas that P4P supporters have most hoped would benefit from this new form of payment is primary care. (15) Fee-for-service reimbursement has traditionally disadvantaged primary care by overpaying for procedures and intensity of care, (16) while underpaying for evaluation and management services that require physicians to spend time diagnosing and coordinating patients' care. (17) This underpayment has led many primary care physicians to feel like hamsters on a treadmill, seeing more and more patients to make up for reimbursements that do not keep up with their practice expenses. …
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.016 | 0.023 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.004 | 0.003 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".