Challenging the supremacy of evidence-based medicine through artificial intelligence: the time has come for a change of paradigms
Bibliographic record
Abstract
Three years ago, at the Davos Summit, Klaus Schwabe, the President of the World Economic Forum, made the following statements: ‘We stand on the brink of a technological revolution that will fundamentally alter the way we live, work, and relate to one another. In its scale, scope, and complexity, the transformation will be unlike anything humankind has experienced before’ [1]. The development of the steam industry at the end of the 18th century is considered to be at the heart of the First Industrial Revolution, while the emergence of electricity and assembly lines (at the beginning of the 20th century) marked the Second Industrial Revolution, and the implementation of computer technology, nuclear energy and mass production in the 1960s represented the Third Industrial Revolution [2]. The current Fourth ‘Industrial’ Revolution (even if it does not impact on industry only), which began in 2010, is characterized by ‘the underlying digital logic that changes everything’ [3] and includes a number of technological developments that the majority of health care professionals consider to be experimental and futuristic: nanotechnology, artificial intelligence (AI), deep learning algorithms, Internet of Things and three-dimensional printing. In reality, neither the European Society of Cardiology nor the Kidney Disease: Improving Global Outcomes Guidelines have included the use of the above technologies in their guidelines. In fact, excellence in health care resides in evidence-based medicine (EBM) [4]. The concept of EBM has become well known worldwide since the 1990s [5]. According to its definition, EBM is characterized by the integration of clinical expertise, the patient’s values and the best available evidence in the process of decision-making related to the patient’s health care. The best scientific evidence is considered to be data derived from multiple randomized controlled trials (RCTs) or meta-analyses, conducted at a population level, that can prove the effectiveness of a particular (new) drug/intervention, as well as the harm and inefficacy of others, in comparison with the best existing therapy [6]. Nowadays, all European medical societies (e.g. those of nephrology and cardiology) endorse an increasing number of clinical guidelines edited by specialist experts, which are continuously updated with new data from medical databases and with the latest objective information from the literature [7]. All of these guidelines begin with the presentation of two tables showing classes of recommendations and levels of evidence, respectively [8]. Moreover, these guidelines are included in the teaching curricula of all medical schools and university hospitals, and all medical students and medical trainees are assessed on their knowledge of these recommendations and whether they act/think according to the best evidence-based guidelines. However, after almost 30 years of EBM supremacy, health care practitioners are finding that many (and increasingly) clinical situations are without clear guidance or ‘best evidence’, or are guided based on various experts’ opinions. One example is that the majority of RCTs on the use of antithrombotic medication (anticoagulant and antiplatelet therapy) in acute coronary syndromes or post-percutaneous coronary interventions (PCIs) or in atrial fibrillation, or both, exclude advanced chronic kidney disease (CKD) patients [9]. Moreover, all the risk scores (for bleeding or ischaemic events) are not validated in the G5D CKD population, thus resulting in a lack of medical evidence-based foundation for specific recommendations [10, 11]. In other words, in the 21st century, antithrombotic treatment (with antiplatelets and/or anticoagulants) of the large number of advanced CKD patients with cardiovascular comorbidities is based on observational studies or expert position statements [10]. It seems obvious that many of the highly awaited RCTs will be delayed for many years due to ethical or logistical reasons. On the other hand, by using various data sources (experimental, environmental, clinical, biological or wearable devices), the process of medical diagnostic or therapeutic decision-making could, and would, be possible thanks to the use of specific ‘machine learning algorithms’ (ensemble decision trees, support vector machines, neural networks, deep learning and topological data analysis) [12]. Thus, in our specific example described above, physicians will be provided with new bleeding risk algorithms based on: (i) ‘big data’ [13] (a term describing a massive volume of both structured and unstructured data—an assemblage of vast information that otherwise would be difficult to process using traditional database and software techniques [14], the main characteristics of which are: volume, velocity, variety, variability and complexity [15]); and (ii) decision pathways generated by AI algorithms (not by RCT-derived evidence). In fact, the answers from such algorithms will be ‘truly individualized’ for every single patient, in a way that is not possible using conventional guideline recommendations. For example, current cardiology and nephrology guidelines are unable to provide evidence (or even advice) on how to treat an 80-year-old lady with G5D CKD, chronic atrial fibrillation and a recent primary PCI (since the bleeding/thrombotic risk scores do not apply in this context and existing antithrombotics have no study-derived evidence for the ‘triple association’ in this age group with advanced kidney disease on dialysis) [9]. In contrast, in this very specific context, a complex AI algorithm will find patterns of bleeding and predict (in a new and ‘deep’ manner) whether this lady will manifest, or not, major adverse events with specific drug combinations. In other words, the medical community will be offered not only (limited) ‘RCT-derived evidence’, but also AI tools with improved (and self-improving) accuracy and predictive power (Figure 1). Scientific interest will shift focus from providing evidence to developing self-learning and pattern-detecting AI solutions. The modern revolution of medicine. Only a few months ago, a first step was made in predicting renal function simply through routine kidney ultrasonography, using deep learning algorithms based on convolutional neural networks [16]. As the authors stated, the algorithms help to optimize cost-effective CKD screening methods without laboratory testing, particularly in settings with limited health care resources. This study changes the way we perceive the relationship between the structure and function of the kidney and ‘demonstrates the possible role of AI in turning conventional images into functional screening and diagnostic tools – this type of automation will be pervasive in the era of AI and Big Data’ [16]. A recent study using AI (Bayesian networks and support vector machines) identified a simple, and yet useful, way to predict the evolution of CKD (especially in early G1 and G2 stages), relying solely on routine data gathered from a general practitioner’s consultation. Thus, the AI tool could predict aggressive evolution of kidney disease in otherwise low-risk CKD subjects [17], in a manner not possible with conventional diagnostic/prognostic guidelines. This is extremely relevant, considering the new strategic directions of the ERA-EDTA supporting research initiatives based on big data derived from registries and large trial databases. Another important example is the recent development of AI predictive models (through machine learning) to facilitate the judicious allocation of kidneys before transplantation to avoid rejection [18]. The model analysed 48 variables (from both pre-transplant donors and recipients) using a machine learning algorithm (based on the principle of Bayesian Belief Network) and was able to predict graft failure within the first year or within 3 years post-transplantation. In a study published this year, which analysed variables from the European Clinical Database (EuCliD) [19], an AI model was developed (using a patient- and session-specific artificial neural network) to predict blood pressure and heart rate profiles, post-dialysis body weight and the Kt/V for each haemodialysis session. The authors claimed that this AI algorithm could accurately predict hypotensive events, which will enable clinicians to balance intradialytic fluid removal (through treatment duration and ultrafiltration rate). Moreover, according to the authors, ‘the case could be further explored by assessing the impact of manipulating dialysate electrolyte compositions or the temperature on heart rate and [blood pressure] changes so that multiple endpoints can be concurrently optimized’ [20]. Such an approach could solve one of the most vexing and subjective treatment situations in nephrology. All the above examples, though very recently published, demonstrate a new way (and a novel paradigm) in which the diagnostic and treatment approaches to CKD can be optimized, thus inviting clinicians and researchers to envision, and expect more from, the use of AI. It is becoming clearer that, in the near future, RCTs will no longer be needed and the paradigm will change—the medical community will be offered self-improving decision tools, based not on RCTs, but on deep learning algorithms. Although fascinating, AI has not been widely validated yet in the medical field. At present, the use of AI in medicine is still a subject of debate (including the issues of various algorithm choices and ethical challenges). In certain medical situations, it would seem wiser to still use a regression logistic model, rather than a machine learning technique [21, 22]. Furthermore, while it was taking a few years to solve an ethical problem such as the Google DeepMind project on acute kidney injury prediction in collaboration with the Royal Free National Health Service Foundation Trust, this research question itself was solved through the application of AI algorithms using different data sets taken from the US Department of Veterans Affairs (published in Nature in August 2019 [23]). In fact, nowadays, the development of AI for use in health care is not as widespread as its application in the retail industry, finance and banking, or the use of intelligent voice assistants (including Apple’s Siri, Google Home/Translate, Amazon’s Alexa or Microsoft’s Cortana). The present of the non-medical domains represents the future in medicine. Currently, these decision-making algorithms (developed on AI platforms) are not (necessarily) designed to replace EBM. In fact, they could be the modern answers to the limitations of contemporary EBM. We believe that this very step is marking the transition from the ‘EBM paradigm’ to the ‘deep medicine concept’ coined by Eric Topol [24]. Furthermore, machines could be trained to learn much about patients, such that both diagnostic and therapeutic recommendations (through neural network algorithms) could be generated with unbeatable accuracy. Clinical practice is taking on a ‘new look’ that simply reflects the implementation of developments from the Fourth Industrial Revolution in the health care industry. Since EBM seems to be representative of the Third Industrial Revolution, one could consider ‘deep learning medicine’ to be the herald of the Fourth Industrial (AI) Revolution. For now, we can envision the evolution of AI medicine as an extension of EBM (providing solutions to unsolved current problems, as discussed earlier), even if Schwab does not agree with the continuity of paradigms (‘There are three reasons why today’s transformations represent not merely a prolongation of the Third Industrial Revolution but rather the arrival of a Fourth and distinct one: velocity, scope, and systems impact’[1]). Finally, even if the current guidelines (based on the latest EBM recommendations) do not include the implementation of ‘deep medicine’ strategies in the decision-making process, in the years to come, we will probably witness a sweeping profound revolution of the current paradigms (that will change the face of medicine as we perceive it today). Maybe a complex think tank comprising academic researchers coming from medical and computer science backgrounds, along with continuous input from industry, will reshape current guidelines by incorporating AI developmental updates. None declared.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.143 | 0.152 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.005 | 0.002 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.006 | 0.060 |
| Scholarly communication | 0.028 | 0.054 |
| Open science | 0.006 | 0.013 |
| Research integrity | 0.025 | 0.059 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".