Bibliographic record
Abstract
Evolve, the Royal Australasian College of Physicians' (RACP) equivalent of the American Board of Internal Medicine (ABIM) Foundation's Choosing Wisely campaign, was launched in 2015. The aim of these campaigns is to demonstrate to the community the profession's tangible commitment to its social justice principles,1 which include effecting responsible stewardship of entrusted resources. Unfortunately, early indications are that the RACP programme is taking the same pathway to failure as the initial phases of the ABIM campaign suggesting the need for a reset. As a percentage of gross domestic product (GDP), inflation-adjusted expenditure for health care in Australia, New Zealand and the United Kingdom has continuously increased for 20 years.2 In real terms, for the period 2009–2014, the average annual growth in per capita health spending in Australia was 2.0% (down from 2.9% in the preceding 5-year period) while in New Zealand, it was 0.6% (down from 4.1%). As a percentage of GDP, in 2014, Australia, New Zealand and the United Kingdom spent 9.4, 11.0 and 9.1% of their respective GDP on healthcare compared to 20 years earlier when each spent circa 7%. In New Zealand, there was a sharp rise from 8.4% in 2007 to 10.7% in 2008. The ABIM Foundation launched the Choosing Wisely campaign in 20123 encouraging specialty groups to adopt the ‘Top Five’ list approach. This approach encouraged the identification of five diagnostic tests or treatments that are amongst the most expensive and commonly ordered by members of a specific specialty but which have been shown by currently available evidence to provide no meaningful benefit to significant categories of patients.4 Over 12 OECD countries have adopted similar programmes. The primary goal is to reduce waste in the healthcare system and avoid risks associated with unnecessary treatment. At the second RACP sponsored Evolve meeting in Sydney on 7 April 2016, several medical specialty societies adopted the choosing wisely approach and presented their lists of low-value tests or procedures. Several items were duplications of ABIM counterparts. There is much interest in the effectiveness of the ABIM Foundation campaign as it is the longest running. To maintain support, such programmes need to demonstrate both short- and long-term effectiveness. Analysis of the short-term impact on medical and pharmacy claims to one USA health insurer covering 25 million members suggests that the approach of developing and promoting lists of low-value tests and procedures has effected little if any change in physician behaviour.5 Do not routinely repeat dual X-ray absorptiometry (DXA) scans more often than once every 2 years. The optimal interval for repeat DXA scans, especially in high-risk patients such as those initiating glucocorticoids, is uncertain. Despite this uncertainty, availability of physician office-based DXA scanners and a Medicare policy that reimburses for repeat DXA scans every 2 years likely contributed to higher utilisation of this technology in some groups of patients. further research is needed on the timing and frequency of DXA testing in patients at risk for osteoporosis and those with known osteoporosis who are receiving treatment. do not use antihistamines to treat anaphylaxis do not order ANA testing without symptoms and/or signs suggestive of systemic rheumatic disease. These are recommendations directed at groups outside the craft group from which they were developed. They are educational statements covered in textbooks, pathology test manuals, test report comments and similar resources. They do not engender ownership in those to whom they are directed and as such, if read at all, are unlikely to have any lasting impact. An approach with potentially greater short- and long-term impact would be to undertake critical appraisal of clinical/practice guidelines. Such guidelines are becoming increasingly embedded in electronic pathways for general practice.7-9 Local medical specialty groups are major contributors to these guidelines, but critical appraisal using standardised instruments is either not done or not acknowledged as being done. In addition, despite their potential impact on consumption of healthcare resources, conflicts of interest of those generating the guidelines are rarely if ever transparently addressed and what if any debiasing strategies are used in development are not apparent. A recent review of management recommendations for osteoporosis in 78 clinical guidelines lodged at the Agency for Health Research and National Guideline Clearinghouse noted that 90% recommended BMD measurement as a monitoring procedure.10 Local guidelines understandably incorporate the same recommendations.7 It is difficult to determine if these guidelines have been evaluated against either international or national validated and endorsed appraisal instruments such as the Canadian Institutes of Health Research's Appraisal of Guidelines for Research and Evaluation (AGREE) II11 or the Australian National Health and Medical Research Council standard for clinical guidelines.12 BMD is one of several predictors of fragility fracture risk. Its utility varies in different geographic populations and ethnicities. In New Zealand, a BMD measurement is one criterion amongst others enabling qualification for a public subsidy for bisphosphonate therapy. However, BMD adds only a small percentage to the predictive power of FRAX, a commonly used online Fracture Risk Assessment Tool developed under the auspices of the World Health Organization.13 Take a 65-year-old woman with one previous fragility fracture and no other risk factors except for a femoral neck T-score of −2.5. Her FRAX calculated risk for a major osteoporotic fracture or hip fracture in 10 years is 17 and 3.6%, respectively, whereas if the T-score is not included it is 19 and 4.9%. In a real world setting, a Canadian study examined the predictive power for fragility fractures of FRAX with and FRAX without femoral neck BMD in a cohort of over 35 000 women aged 50 years and over.14 The cohort was divided into those not treated with anti-resorptive therapy, those currently treated and those treated in the past. The aim of the study was to determine how the predictive power of FRAX performed for fragility fracture in these groups. Of relevance to the current discussion was the finding that the addition of BMD to the FRAX score was not additive to the predictive power of FRAX without the inclusion of BMD. Nevertheless, irrespective of the comparatively small weighting BMD gives to predicting fragility fracture, reduced baseline BMD does increase the risk. This observation and the incorporation of BMD monitoring into guidelines has cemented its use into routine medical practice. Treatment with bisphosphonates over 3–5 years has been shown in multiple large clinical trials to be associated with an increase in BMD and an absolute risk reduction in fragility fractures in post-menopausal women at significant risk: vertebral fractures 5% reduction, hip fractures 1% reduction and other fractures 2% reduction.15, 16 Despite strong objections, this association between increases in BMD density and absolute risk reduction in fracture in mainly post-menopausal osteoporosis has been used to justify the use of BMD as a surrogate end-point for fragility fracture in osteoporosis treatment trials in other groups at risk. Absolute reduction in fragility fractures from bisphosphonate therapy in those taking corticosteroids is not defined although BMD in these mainly rheumatic disease patient populations is either maintained or increased after 3–5 years of therapy.17 In other meta-analyses of treatment effect in steroid-associated osteoporosis, there was a statistically significant relative risk reduction in combined asymptomatic and symptomatic vertebral fracture but not in hip fractures.18 The optimal duration of bisphosphonate therapy is undefined.19 Reports of bone complications such as the increased risk of ‘atypical’ femoral fractures sounded a cautionary note to the long-term safety of these agents. Given the mechanism of action of bisphosphonates, this not unexpected complication prompted the recommendation that after 5 years of therapy patients should be reviewed as to the appropriateness of continuing therapy.19 What predictors of future fracture risk should be used to guide therapeutic interventions during and at completion of anti-resorptive therapy is unknown but several studies have addressed this question with respect to BMD. In 2005, Watts et al. combined the data of three phase III randomised, double-blind placebo-controlled trials of post-menopausal women with osteoporosis treated for up to 3 years with risedronate (n = 2561) where fracture was the trial end-point.20 They aimed to study the relationship between treatment-related change in BMD and reduction in the incidence of non-vertebral fracture. They concluded that BMD at neither the spine nor the femoral neck as measured by DXA predicted the degree of reduction in non-vertebral fractures. The Fracture Intervention Trial long-term extension (FLEX) trial21 examined at-risk post-menopausal women who had received 4–5 years of alendronate therapy. Participants were randomly assigned to either receive placebo or continue alendronate therapy for a further 5 years. The FLEX trial did not show any statistical difference in the absolute risk of hip fracture or other non-vertebral fractures between the placebo and treatment groups. There was a 2% absolute reduction in the number of clinical vertebral fractures in the treatment group. In another study with a similar extension phase (HORIZON-PFT)22 but using a different bisphosphonate, there was a 2.5% reduction in morphometric (asymptomatic and radiographic) vertebral fractures and no demonstrable reduction in hip and other fractures. Further analysis of pooled data from these extension trials suggested that there may be no benefit in reducing vertebral fracture risk or reducing non-vertebral fracture risk from continuing bisphosphonate therapy beyond 3–5 years.19 The FLEX study also examined the clinical utility of repeat BMD measurements and markers of bone turnover in identifying subjects at risk following alendronate therapy.23 While BMD showed a moderate decline in those discontinuing alendronate therapy and a gradual rise in biochemical markers of bone turnover neither of these measures was useful in prediction models of bone loss in those continuing alendronate or those switched to placebo. This poor predictive power of BMD monitoring has been noted for over 10 years.24, 25 It is difficult to explain why it continues to be recommended as a tool for monitoring fracture risk in treated patients or those stopping therapy without invoking cognitive dissonance: that resistance to act on evidence that runs counter to deeply embedded beliefs that are reinforced by routine practice. Our understanding of the influence of cognitive bias in decision-making owes much to the pioneering work of Amos Tversky and Daniel Kahneman.26 Over 40 years of research has revealed reproducible patterns that are generic to both professional and day-to-day life. No individual, no matter how intelligent, is immune. However, recognising their existence and the scenarios in which they are prone to operate may alert clinicians to enact debiasing strategies in decision-making processes. For ease of memory, a check list that uses a humorous intuitive classification of the more insidious and dangerous cognitive biases called the Seven Deadly Sins (Table 1)27 is one tool that may be used at times of critical decision-making. If such a check list or other debiasing strategy is used along with the use of validation tools in the development of or review of osteoporosis and other guidelines a priori, it is likely that recommendations lacking adequate evidence, such as the inclusion of BMD for monitoring treatment effectiveness would not have been included. However, if we take the history of BMD monitoring as a guide, such measures may not be enough to overcome the power of cognitive dissonance. Collaboration with medical jurisdictions and unbiased stakeholders is likely to be required in order to dismantle the dogma of established practice. New Zealand Rheumatology Association (NZRA) for endorsing the author as its representative at the RACP Evolve meeting 2016. The RACP is acknowledged for sponsoring the Evolve campaign workshops. The opinions expressed in this article are those of the author. They are not necessarily endorsed by the NZRA, the RACP or any other body.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.021 | 0.171 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.002 | 0.012 |
| Insufficient payload (model declined to judge) | 0.007 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".