MétaCan
Menu
Back to cohort
Record W4407841731 · doi:10.1113/ep092157

The state of mechanistic research in the evidence‐based medicine era: A sandwalk between triangulation and hierarchies

2025· editorial· en· W4407841731 on OpenAlexfundno aff
Ronan M. G. Berg, Cody Durrer, Jan Kyrre Berg Olsen Friis, Mathias Ried‐Larsen

Bibliographic record

VenueExperimental Physiology · 2025
Typeeditorial
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsnot available
FundersCanadian Institutes of Health ResearchTrygFonden
KeywordsTriangulationState (computer science)MedicinePsychologyData scienceComputer scienceGeographyCartographyAlgorithm

Abstract

fetched live from OpenAlex

Formally, to have evidence is to have a ‘conceptual warrant for belief or action’ (Goldenberg, 2006), and to paraphrase Frank Herbert (1926–1980) in his novel Dune, evidence is certainly the melange, or ‘spice,’ around which all science is centred (Herbert, 1965). The concept of evidence that seems to dominate most biomedical sciences these days is that advocated by the evidence-based medicine (EBM) movement, which emerged as a new paradigm in the early 1990s with the ambition of basing clinical practice and teaching strictly on evidence as defined within the ‘hierarchy of evidence’ (Timmermans & Mauck, 2005) (Figure 1). Despite our enduring quest for truth within the field of experimental physiology, which seems more important than ever these days (Drummond & Tipton, 2024), mechanistic studies are consistently placed at the bottom in various incarnations of this hierarchy (Djulbegovic & Guyatt, 2017; Williamson, 2019). Does this mean that all our efforts as researchers within experimental physiology and other lines of mechanistic research contribute nothing more than ‘low quality’ evidence? Certainly not! Here, we will make the case that while EBM and experimental physiology benefit from each other and are complementary in many ways, they operate with fundamentally different frameworks. The main scientific philosophical concepts we will discuss are summarised in Box 1. We will make the case that EBM-based criteria for what constitutes ‘good evidence’ cannot uncritically be extrapolated to mechanistic research. And vice versa for that matter. Empiricism: The idea that knowledge comes from sensory experience and observation. It emphasises that all knowledge must be grounded in empirical evidence, meaning it must be rooted in observable phenomena. Empiricism rejects the notion that innate ideas or purely logical deduction can provide meaningful knowledge of the world. Instrumentalism: A view often linked to pragmatism; it holds that theories are valuable only insofar as they are useful in solving problems. Theories are seen as tools, not necessarily reflective of the true nature of the world, but as instruments for achieving practical outcomes. Positivism: An extension of empiricism, this philosophy, which historically exists in several incarnations, asserts that knowledge of the world should be based strictly on quantifiable data from observation and experimentation. It maintains that empirical data are the only valid sources of knowledge, rejecting metaphysical explanations or speculation as unscientific. Positivism is closely related to scientific materialism, as both reject non-material explanations and emphasise empirical investigation, but positivism focuses specifically on observable and measurable phenomena. Pragmatism: A practical approach that emphasises knowledge as a tool for solving real-world problems, particularly in settings like clinical medicine. Pragmatism values the effectiveness of interventions over theoretical completeness or certainty, focusing on what works in practice, even if the underlying theory is uncertain or incomplete. Scientific realism: The view that a real world, including both observable and unobservable entities, exists independently of our perceptions or theories. Scientific realism holds that scientific theories aim to describe and explain the true nature of this world, and when a theory is successful, it is likely because it accurately reflects reality. Scientific realism aligns with scientific materialism in its commitment to understanding the world as independent of human perception, though it also accepts that unobservable entities, such as subatomic particles, are real and can be studied through scientific inquiry. Scientific materialism: A philosophy closely aligned with positivism and often linked to scientific realism. It holds that all phenomena, including mental processes and consciousness, can be fully explained by physical matter and its interactions. Like positivism, scientific materialism emphasises empirical investigation as the primary means of understanding the world, rejecting non-material explanations as unscientific. However, while positivism focuses strictly on observable and quantifiable data, scientific materialism extends this to assert that even unobservable processes, such as those studied in physics and biology, must ultimately be rooted in material causes. This places it within the broader framework of scientific realism, as both posit that a real world exists independently of human perception and that it can be known through science. The field of experimental physiology emerged in France and Germany during the early 19th century, liberating physiology from natural philosophy, which had dominated its early history (Bailey et al., 2023a; Cunningham, 2002). Consequently, scientists in this new field radically discarded the concept of vitalism, which posited a mystical life force as the distinguishing element between living and non-living matter. Instead, life's processes were to be determined within the realms of scientific materialism, that is, by the laws of physics and chemistry alone and thus requiring analysis through these exact scientific disciplines (Coleman, 1987; Culotta, 1970). While this new field of experimental physiology was a basic science, its scope was quite practical: to inform clinical practice by providing physicians with a theoretical framework upon which to base their decisions (Goldenberg, 2006). At its core, experimental physiology is a positivistic science, founded in so-called scientific realism, which is based on the premise that a real world exists, with its own structural and functional properties, which can be objectively studied through experimentation (Goldenberg, 2006). Although it underwent gradual modification, this positivism, with its reliance on deductive and inductive reasoning from mechanistic principles and theories, dominated clinical medicine for most of the 20th century (Timmermans & Mauck, 2005). EBM specifically evolved as a reaction against this, as mechanistic studies and theories were repeatedly found to be ineffective in predicting treatment effects in the clinical setting, and sometimes even leading to practices that turned out to be directly harmful. Even from the standpoint of scientific realism, this makes sense, because experimental models and conditions rarely reflect clinical reality in all its complexity, such that seeking for monocausal relationships for phenomena that are in reality multicausal has a high likelihood of failure. Furthermore, when one mechanism is targeted, its operation can be stopped or screened off by another causal factor, because several mechanisms typically operate simultaneously in vivo. Only the joint knowledge of all mechanisms and their interactions would suffice for taking mechanistic causal claims as a basis for decisions regarding interventions, something that even the most skilled experimentalist would find extremely challenging, if possible at all! EBM advocated for a radically different approach. Rather than relying on mechanistic theories, treatment effects should be evaluated specifically on the basis of clinical and patient-centred outcomes. These include mortality, hospital discharge or readmission, changes in treatments and health-related quality of life, as well as through the systematic registration of adverse events (Timmermans & Mauck, 2005) and/or by changes in context-specific biomarkers that are known to correlate with a given clinical outcome of interest (Manyara et al., 2024). While the exact philosophical underpinnings of EBM are unclear and a matter of debate (Djulbegovic & Guyatt, 2017; Kulkarni, 2005; Thomas, 2023), it is undeniably positivistic, yet clearly opposes the scientific realism of mechanistic reasoning, at least in the context of clinical decision-making. Rather, EBM has clear elements of pragmatism, that is, emphasising the practical application of knowledge over theoretical reasoning, such that knowledge comes from observations and experiences rather than innate ideas or reason. Taken to the extreme, this implies that it is only relevant if a given cause–effect relationship is present or not, and any theoretical considerations of how and why are irrelevant (Goldenberg, 2006). While the purpose is thus not theory-building with the goal of understanding the real world, EBM mainly has an instrumentalistic approach to theories. This implies that EBM may accept theory-building as a tool for predicting and controlling phenomena, but without considering such theories true descriptions of the real world. This is consistent with the fact that studies conducted within an EBM framework often test hypotheses regarding treatment effects based on mechanistic theories. Indeed, it would rarely be considered ethical to conduct a clinical trial on a new treatment with no theoretical rationale to support its benefit. When it comes to obtaining evidence to inform clinical decision-making, EBM requires that the efficacy and adverse events of a treatment are systematically evaluated on patients in the clinical setting. The primary focus is on design to minimise confounding variables; this is principally achieved through randomised controlled trials (Williamson, 2019). However, it is important to note that considerations regarding what constitutes ‘good evidence’, as depicted in the hierarchy of evidence, specifically relate to the ability to inform clinical decision-making. Thus, the placement of randomised trials at the top and mechanistic studies at the bottom does not reflect a difference between their implicit scientific value, but merely in their utility for directly informing clinical decision-making and health recommendations. Hence, just as mechanistic studies are insufficient for informing clinical decision-making, clinical studies—here understood as studies on patient populations focusing on clinical and patient-centred outcomes—are rarely suited for making causal claims regarding mechanisms (Maziarz, 2023). Here, it may be worth noting that due to their different philosophical underpinnings, EBM and mechanistic research operate with different concepts of causality. As EBM has its philosophical roots in pragmatism, it builds on a strictly manipulative causality concept, asserting that the experimental manipulation of a cause will result in the manipulation of an effect, thus practically making randomised controlled trials the only means for estimating the average treatment effect and potential harms, provided that they are well designed and well conducted. As such, the best evidence is obtained when the same hypothesis regarding a treatment effect repeatedly resists falsification in similarly conducted studies in similar populations. In contrast, causality is best appreciated as pluralistic, relying on both manipulative and descriptive reasoning in mechanistic research (Maziarz, 2023; Williamson, 2019). While causal claims can strictly be applied only to the specific experimental setting, model and system under study in mechanistic research, scientific realism permits interpretation of cause–effect relationships within its own adaptable theoretical framework, provided they can be replicated and consistently resist experimental falsification. This collective evidence is then used to make general claims about physiological mechanisms. Indeed, it is this ‘inductive leap’ of scientific realism that builds theories from which new hypotheses can be formulated, including clinical studies conducted within an EBM framework, such that these do not rely solely on incidental discoveries to generate new ideas for potential therapies. While the use of mechanistic studies to inform working hypotheses for clinical trials is the modus operandi, it works both ways. Clinical trial results can also yield incidental findings that inspire new mechanistic hypotheses, prompting further research. A notable example is the Women's Health Initiative trial, which unexpectedly revealed that combined oestrogen–progestin hormone replacement therapy increased the risk of breast cancer and cardiovascular disease compared to placebo in post-menopausal women (Rossouw et al., 2002). This was contrary to the prevailing belief that hormone replacement therapy might be protective against these conditions and given that the Women's Health Initiative trial did not provide the actual mechanisms, the findings led to a long line of mechanistic studies into the roles of oestrogen and progestins in both carcinogenesis and vascular function. Similarly, a randomised controlled trial of intensive insulin treatment to maintain strict blood glucose control in critically ill surgical patients showed that this reduced mortality specifically due to septic complications (Van den Berghe et al., 2001), which led to subsequent mechanistic studies on the immune-modulatory effects of insulin and related peptides. Mechanistic research has learned many important lessons from EBM, particularly the emphasis on design, where random assignment to treatment and control groups ensures, at least in principle, an equal distribution of unknown confounders, although these may be unequally distributed by chance, particularly if the sample size is small. This is notably important in clinical trials because the risk of unknown confounders that are unbalanced between groups is particularly high for clinical populations. Experimental physiology and related fields within mechanistic research, such as pharmacology, biochemistry and microbiology, encompass a wide range of research methods in both laboratory and applied (i.e., environmental or clinical) settings. For experimental physiology, this includes in vitro, ex vivo and in vivo models, which are applied to a variety of systems, including individual cells, cells in culture, tissue preparations and various isolated organ preparations, as well as animals and studies involving human subjects. From its inception, experimental physiology drew from the controlled experiments that formed the basis of physics and chemistry (Coleman, 1987). In its simplest form, the controlled experiment involves predicting an event by assessing the impact of changes in preconditions within a highly controlled environment. The experimental conditions are then systematically modified and adapted to manipulate or observe spontaneous changes in an independent variable while standardising conditions to rule out the effects of other confounding variables. This is done as part of an iterative process that often involves multiple cycles of switching between deduction and induction, thereby identifying cause–effect relationships between the independent and dependent variables. While successful randomisation is in principle the only means of eliminating both known and unknown confounders, several other procedures may be used to effectively minimise this in mechanistic studies. This includes various experimental manipulations that target the biological pathways under study through pharmacological, environmental, behavioural and/or genetic activation or inhibition. This may be relevant when randomisation is either impossible or unethical, such as when the natural history of a disease is studied or when a disease is compared to the healthy state, or when the exposure under study is assumed to have harmful effects, as based on theoretical reasoning or other lines of empirical evidence. Furthermore, controlling for various known confounders in the statistical analysis can also be effective here, for example, via inclusion of covariates in the analysis or by weighted regression. Despite some thematic overlap, it is important to note that although experimental physiology has historically been conceived to inform clinical medicine, it is a basic science with various aspects of physiology having a much broader scope than health-related outcomes. Challenges may thus arise when mechanistic studies in humans are classified as clinical studies for legal or ethical reasons, often leading to the mistaken belief that this classification also applies to the scientific aspects of the study (Richter et al., 2024). Clearly, when conducting studies on humans in various applied (including the clinical) settings, the risk of confounders is higher than in the controlled laboratory setting—for we can only to a very limited extent control previous or concurrent factors, including various genetic and environmental exposures, that may contribute to the observed cause–effect relationships. Of note, the same applies to studies on non-laboratory animals, such as pets, livestock and wild animals. While randomisation is probably the most powerful tool for avoiding unbalanced unknown confounders between groups, it is important to note that randomisation in itself does not control for confounders if improperly implemented or if sample sizes are too small. Along with imprecise outcome assessments, and poorly described experimental set-ups, this increases the risk of so-called magnitude and sign errors (Gelman & Carlin, 2014). Furthermore, properly performed randomisation with adequate sample sizes is not always possible or even relevant in mechanistic studies, particularly in the applied setting. In any event, it is useful to clearly distinguish actual hypothesis-testing studies from exploratory studies. We posit that many different designs can be used to test mechanistic hypotheses in applied settings, but the highest degree of is achieved when based on a randomised controlled design, in which it is possible to and an effect that is, a relevant for sample size Consequently, most studies in the applied are as they rarely focus on a variable or outcome but rather on several that a of the of Furthermore, it is often impossible to the difference that is However, as we will results from exploratory studies are by no means to be considered ‘low quality evidence’ as in to a physiological difference that is relevant is a matter of with some emphasising that a relevant difference should be determined by either theoretical or statistical or & et al., 2023), and that it should be at the relevant difference et al., 2023). However, the is rarely for the given context and only reflects the utility of the variable as a of a given clinical and thus not its mechanistic one that any measurable may in principle be thus the of the and of the specific physiological at et al., 2023). However, one has it often impossible to a hypothesis based on this, making statistical a in many physiological studies et al., 2024). A of EBM in its ability to systematic and which findings from multiple studies and provide an of the degree of the of evidence regarding treatment effects on in clinical populations. While mechanistic evidence on collective evidence from different lines of research rather than individual studies, it can benefit from a similar approach to data through systematic and to the degree of of evidence in of a given Rather than focusing on randomisation as within the evidence this should be based on so-called Here, involves the use of multiple models, and settings to one each with its own and because results that these different models and are likely to be & This approach would also that findings from many mechanistic studies, both hypothesis-testing and including the many individual studies that are due to of and ethical considerations et al., & all contribute to the collective evidence on & It is clear that both EBM and mechanistic research to evidence in the of a warrant for belief and but from quite different philosophical However, as the within all life science research has the of the (Bailey et al., it is relevant to discuss and what constitutes ‘good evidence’ in experimental physiology and other of mechanistic research. This should be done while in that to the principles of EBM would more than of research within our it is practice to what is or evidence based on However, just as our field has from the of study and analysis to analysis and et al., 2024), an approach also by the in this et al., 2024), may it be to for data and in systematic and on mechanistic Frank Herbert in to describe the so-called but with a and This for the of the is used to the and that the Similarly, for data and in systematic and of mechanistic studies will be a between the different and the thematic overlap, mechanistic research fundamentally from clinical research, focusing on causal claims regarding mechanisms rather than informing clinical decision-making. This that rather than merely the best of In our for data and in systematic and in mechanistic research would a approach that the basis for criteria to the of evidence. While the exact of to be we the model in evidence a mechanism is dependent findings that are and by multiple sources of data, that is, models, and/or multiple and multiple experimental any about a mechanism should to the system and in which they were In the more data methods and experimental manipulations that a the the of evidence. this approach to the of for evidence and a of its certainty, to the of and framework used in EBM & 2024). As Frank Herbert (Herbert, the the While evidence may be the of any scientific this is too for any when considering how it is in reality to control even the simplest in the laboratory or applied setting. The systematic of the results are and they can be replicated and may that the ‘inductive leap’ to can be the of the mechanism into physiological The must have and the of this and to be for all aspects of the in that related to the or of any part of the are and as for and all those for are is by had no in the to or the other have any of interest to The for is by and The had no in study design, data and to or of the was by the of Health The had no in the to or the

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.115
metaresearch head score (Gemma)0.052
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.732
Threshold uncertainty score0.956

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.1150.052
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0030.000
Bibliometrics0.0010.001
Science and technology studies0.0000.001
Scholarly communication0.0000.000
Open science0.0020.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.737
GPT teacher head0.615
Teacher spread0.122 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations10
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueExperimental PhysiologySame topicMeta-analysis and systematic reviewsFrench-language works237,207