MétaCan
Menu
Back to cohort
Record W4375850928 · doi:10.1002/mds.29420

A Unified Framework for Evidence‐Based Diagnostic Criteria Programs in Movement Disorders

2023· article· en· W4375850928 on OpenAlexafffund
Tiago Mestre, Margherita Fabbri, Sheng Luo, Glenn T. Stebbins, Christopher G. Goetz, Cristina Sampaio

Bibliographic record

VenueMovement Disorders · 2023
Typearticle
Languageen
FieldMedicine
TopicParkinson's Disease Mechanisms and Treatments
Canadian institutionsOttawa Hospital
FundersDystonia CoalitionGenentechNational Institutes of HealthNeurocrine BiosciencesCanadian Institutes of Health ResearchParkinson CanadaSunovionRush UniversityCleveland Clinic FoundationCleveland ClinicBiogenInternational Parkinson and Movement Disorder SocietyUniversity of OxfordAdamas PharmaceuticalsCHDI FoundationParkinson's FoundationPfizerUniversity of OttawaAlzheimer's AssociationMichael J. Fox Foundation for Parkinson's ResearchU.S. Department of Defense
KeywordsOperationalizationMovement disordersParkinsonismFlexibility (engineering)Diagnostic testMedicineDiseaseAutonomySet (abstract data type)PsychologyMedical physicsComputer sciencePathologyPediatrics

Abstract

fetched live from OpenAlex

The establishment of a diagnosis or definition of an illness1 is central to medicine. Diagnostic criteria correspond to the set of symptoms, signs, and tests that identify individuals with a disease in clinical research and/or clinical practice. Studies adopting validated diagnostic criteria generate knowledge about pathophysiology, prognosis, and therapeutic development. Diagnostic criteria must incorporate the heterogeneity of a disease to identify as many affected individuals as possible and, well anchored in the tradition of Jean-Martin Charcot, must include the archetype, as well as variants and cases of partial expression (formes frustes).2 The development and validation of diagnostic criteria with high accuracy are timely in our field and require both rigor and flexibility. One major challenge is the opposing tension between research and clinical practice, because a clinical practice definition requires widely available diagnostic items, whereas research allows incorporation of innovative, though less accessible, methods with high diagnostic value, which nevertheless must be independently established. The International Parkinson and Movement Disorders Society (MDS) has sponsored projects for developing diagnostic criteria for movement disorders, including several forms of parkinsonism.3-5 Although these high-profile programs have society sponsorship, each has worked under its own autonomy without a centralized framework. In this Viewpoint, we present a proposal of a framework designed to be a uniform, but still flexible, method to develop, validate, and operationalize diagnostic criteria in movement disorders. We start by providing a summary overview of prior projects related to parkinsonism, but the principles will apply to other diagnostic criteria programs with the MDS. Importantly, however, we focus on diagnostic criteria and not on programs related to disease-staging or clinical classification schemes, although some principles and concepts may apply. Neurodegenerative parkinsonisms are the movement disorders with more diagnostic criteria published in recent years, namely Parkinson's disease (PD), multiple system atrophy (MSA), progressive supranuclear palsy (PSP), and dementia with Lewy bodies (DLB). In 2014, a position paper on PD diagnostic criteria highlighted the need for new diagnostic criteria.6 Themes emphasized as requiring revision were the use of neuropathology as a gold standard, the expert clinical examination as the benchmark for diagnostic criteria, descriptive limits of the clinical heterogeneity of PD, and the justification of an identifiable prodromal phase.6 As a result, the MDS Clinical Diagnostic Criteria for PD3 and the MDS Research Criteria for Prodromal PD risk7 were completed, followed by the MDS Criteria for Clinically Established Early PD.8 The MDS-PSP criteria4 and the recently published MDS-MSA criteria5 are examples of other MDS-supported diagnostic criteria projects in other forms of neurodegenerative parkinsonism. In parallel, the DLB Consortium elaborated the revised criteria for DLB9 and prodromal DLB.10 On the surface, these diagnostic criteria programs adopted similar processes such as an initial literature review and a consensus methodology for the generation of a list of diagnostic items that served as the basis for the development of specific diagnostic criteria. However, there was significant methodological heterogeneity across these different diagnostic criteria projects (Tables S1 and S2), so that direct harmonization of approaches is not possible. The successes of these programs with the lack of a uniform diagnostic strategy prompt us to consider the opportunity for a standardized data-driven process for development and reporting of diagnostic criteria in movement disorders. Inspired by the meritorious work developed so far on diagnostic criteria for movement disorders and encouraged by the availability of new methodologies and statistical models, for future diagnostic criteria programs, we advocate a standardized approach applicable across the different movement disorders that could be adopted by the MDS. Such an approach would address methodology and workflow issues for diagnostic definitions within any single disease entity and permit comparisons across diagnostic criteria. In our view, such a format could be applied to new programs, and if successful and well accepted, it has the potential to be utilized or adapted as part of regularly updated versions of former MDS diagnostic criteria projects. Evidence-based practice is the conscientious, explicit, and judicious use of current best information in medicine.11 Though more commonly used to evaluate the efficacy and safety of therapeutic interventions, evidence-based approaches are also applied to evaluate the accuracy and precision of diagnostic tests, including diagnostic criteria.12, 13 Different initiatives have guided the discipline of evidence-based processes. Currently, the Grading of Recommendations, Assessment, Development and Evaluations (GRADE)14 methodology is the most widely applied, with the official endorsement of more than 100 organizations worldwide,15 and it is currently used by the Evidence-Based Medicine in Movement Disorders Committee of the MDS. GRADE is a transparent framework for developing and presenting summaries of evidence and provides a systematic approach for making recommendation,15 which has the quality of combining the standardized appraisal of evidence with the subjective judgment about the certainty of a recommendation. A consensus approach has been a core and standard part of diagnostic criteria projects, including those sponsored by the MDS, aiming to engage relevant stakeholders and deliver a product that is recognized by all interested parties optimizing its peer recognition and uptake by users. Consensus approaches can be broadly divided into informal and formal. There are several limitations to informal consensus methods mostly related with group dynamics and the impact of dominant panel members. The decision-making process of developing diagnostic criteria with the valuation of each candidate criterion for a final diagnostic scheme should adopt a formal consensus method. Different methodologies of formal consensus have been used in the health sciences field with distinct merits and caveats (Table 1). Formal consensus methods can be logistically burdensome, which must be considered when adopting a particular method. We favor a robust critical appraisal of evidence using GRADE combined with the modified NIH Conference consensus method, the least formal of the formal methods to ensure both rigor and flexibility. Each participant expresses his or her opinion freely and impersonally Participation of a large number of participants Affordable No personal contact between experts Depends on a questionnaire design Personal contact between experts Does not allow any individual to dominate discussion (depends on moderator) Certain members of the panel can take over the discussion and drive results. Cost and time Synthesis of literature prior to consensus Multidisciplinary panel encourages consensus from a wider group. Cost and time Requires voting on multiple case scenarios Wide circulation Unbiased panel Interaction is not structured. The used aggregation methodology is implicit—a formal feedback system is lacking. As a model for standardized development program for diagnostic criteria in movement disorders, we propose an explicit and auditable evidence-based evaluation of available data using best-available practices to inform a multistep decision-making formal consensus process involving experts in the field (Fig. 1). The project team would be composed of various groups with the role of project oversight (steering committee), decision of key aspects of the study (decision-making committee), and evidence gathering and appraisal (theme-focused working groups). We envision the participation of content experts and inclusion of all invested stakeholders, including people with the given condition under study and their caregivers. This organizational structure is an example, among others, of how different roles may be allocated to meet the goals of a diagnostic criteria program in movement disorders as currently proposed. After assembling the team for the project, the initial task is to generate a list of candidate diagnostic items that will guide the development of a comprehensive set of diagnostic criteria. We envision the steering committee requesting the theme-focused working groups to present an initial list of questions to a decision-making committee, which through a formal consensus process (eg, modified NIH Conference consensus method as suggested earlier) has iterative discussions until reaching consensus, or unresolved divergence is observed. The definition of an agreement among peers (typically ≥80%) is defined a priori. The list of candidate diagnostic items is generated using a master format of a clinically relevant question. The Population-Intervention-Comparator-Outcome (PICO) paradigm has been used in evidence-based research. We introduce this paradigm to be applied to the current framework. P—defining the target population. The P of PICO stands for population definition, which corresponds to the group of people affected by the disease of interest, in our case, a given movement disorder. In this context, compiling an accurate set of diagnostic criteria requires an agreed-upon and well-defined reference against which diagnostic candidates are measured. Whereas pathology is usually required for the definite diagnosis of a movement disorder, time-honored autopsy criteria are infrequently available4 and exceedingly rare for anchoring diagnoses during life as part of clinical practice. Consequently, for most movement disorders, the team establishing diagnostic criteria may need to consider a highly pragmatic disease definition. For example, in the MDS diagnostic criteria for clinical PD,3 a clinical expert diagnosis was adopted as the gold standard. For prodromal populations, a clinically based definition of the disease is not ideal and biological markers may be crucial, though we recognize that applicable examples of genetic, biochemical, or imaging biomarkers (among others) are still limited in movement disorders. Currently, genetically defined conditions, such as Huntington's disease, in which a universally valid biological biomarker for the disease already exists are the exception. I—intervention (clinical feature or biomarker to be considered in support of diagnosis under consideration). In past diagnostic criteria reviewed by us, the list of candidate items was obtained using a variable combination of data ranging from literature search, pathological data, expert feedback, and the adoption of different consensus methodologies (Table S1). We propose that the clinical features and biomarkers considered for the diagnostic criteria are identified by the theme-focused working groups and complemented with a systematic review (SR). In addition, operational definitions of candidate diagnostic items5 are developed, which are of particular relevance for clinical features (Table S2). C—comparison for the purpose of diagnosis. An appropriate comparison for diagnostic criteria needs to consider its context of use and can vary across different diagnostic criteria programs in movement disorders. The adopted comparison or C can be a group of subjects of interest with a similar but different diagnosis that is clinically pertinent. Healthy nondiseased controls may also be considered as comparators, especially in prodromal states. In all situations, the aim is to select the best test or to create a hierarchy based on diagnostic benchmarks or ease of use that defines a newly defined disease relative to the chosen comparator group. It is also possible to create PICOs to compare two different diagnostic tests, one against the other, in the same population. This type of comparison is relevant to decide which diagnostic items to consider if there is more than one way to measure a characteristic, one being a gold standard or widely accepted diagnostic criteria. O—outcome. The evaluation of the performance of an individual candidate diagnostic item requires the use of a meaningful measurement. Common core diagnostic accuracy measures such as true negative (TN), false negative (FN), false positive (FP), and true positive (TP) are used to quantify diagnostic performance. Other related measures, namely pooled sensitivity (TP/TP + FN), pooled specificity (TN/TN + FP), accuracy (area-under-the-curve analyses), positive likelihood ratio (sensitivity/1 – specificity), and negative likelihood ratio (1 – sensitivity/specificity), have been used. Under the current framework, we propose a hierarchization of these diagnostic accuracy measures based on their relevance and implications for care and research in a patient-centered manner that may differ from project to project. Examples of these implications are the timing of the diagnosis, accessibility to standard of care or novel experimental treatments, direct harm from test or clinical assessment, cost, and impact on patient quality of life or wellness. Using an SR approach, relevant data are collected to answer each clinically relevant question and then synthesized into different levels of recommendation to consider a candidate diagnostic item for final diagnostic criteria. We propose this step in the framework is completed by theme-focused working groups covering domains such as clinical, imaging, wet biomarker, genetics, or pathology. The data collected in the SR may include the analytical validation (reliability and validity) of the quantification process for a candidate diagnostic measure (likely more applicable to a biomarker test) and the clinical validation using the documentation of the diagnostic accuracy measures listed earlier. The level of evidence and its determination follows a semiquantitative approach in which a reviewer takes as starting point an agreed-upon high-quality study design and weighs in a list of factors that downgrade or upgrade the level of evidence,17 namely the risk of bias, imprecision, inconsistency, indirectness, and publication bias (Table 2). The combination of benchmark pooled diagnostic accuracy and a quality score for the evidence (eg, high, moderate, and low) determines a recommendation level for each candidate (Table 3). This appraisal process would consider the quality of evidence regarding diagnostic accuracy and the clinical implications of the adoption of a diagnostic criterion under evaluation. For example, when a candidate diagnostic measure does not enhance diagnostic performance per se, a reviewer may value the ability to provide an earlier diagnosis or its ease of use or accessibility. In our proposed organization, the decision-making committee is responsible for providing the first solution to the composition of the diagnostic criteria items, based on data on diagnostic accuracy and a quality of evidence produced by theme-focused working groups and recommendation level for each candidate, and determining inclusion/exclusion for each candidate diagnostic criterion through an explicit and well-documented process. In cases of revised diagnostic criteria, the committee would consider existing diagnostic items (with changes when necessary) and integrate new items when applicable. A consensus process would be adopted, where all group members participate and vote on the essentiality of each criterion in successive iterations. To develop accurate and feasible diagnostic criteria for movement disorders, statistical modeling such as logistic regression, Bayesian classifier, random forest analysis, and support vector machines may enable the development of the diagnostic criteria by providing information on maximum accuracy and feasibility for different combinations of diagnostic items. Inspiring examples exist, such as the use of logistic regression to develop and validate a prediction model for PD diagnosis based on presentations in primary care23 and of a Bayesian classifier used to update the probability of prodromal PD by integrating diagnostic information that combined estimates of background risk with the results of diagnostic marker testing.7 The resultant and likely novel diagnostic criteria would then be submitted for critique and comment to a larger audience of stakeholders. In the movement disorders field, it may include a larger group of experts, patient organizations, the pharmaceutical industry, possible insurance, and government panels. If organized in the context of MDS, the full MDS membership would be included at this phase. The result of this consultation would be documented in writing and, when deemed appropriate, may lead to updates or refinement of the original conclusions. All the aforementioned steps and deliberations produce a final definition of diagnostic criteria that can be field tested and validated. This process is aimed at testing the criteria to establish the diagnostic performance of each criterion and to determine the combination of items with the best diagnostic performance, that is, the ability to identify as many individuals as possible with a movement disorder with the most parsimonious number of diagnostic items. A regular needs assessment to update diagnostic criteria is always necessary to determine if new data in the form of novel diagnostic tests or knowledge of the disease are mature to the points of needing inclusion or challenging the current diagnostic criteria. As seen with the development of MDS rating scales and other society programs, the process will necessarily be an international one involving all sections of MDS to meet our society mission and achieve a global definition of the given movement disorder under consideration. An explicit data-driven approach provides a transparent, unbiased, and auditable process to reach valid and applicable diagnostic criteria. The proposed framework brings best practices in evidence-based methodology to the diagnosis of movement disorders and will certainly generate rich discussion and comments. Some of the practices embodied in our suggested approach have already been applied in existing MDS-sponsored programs. Our suggestions focus on the proposition that a consistent approach across different diagnoses will allow comparison within and among movement disorders. We are inspired by the experience of other evidence-based programs that have operated under similar principles in the fields of therapeutics and clinical measurement in movement disorders. It is our view that this framework may serve as guidance to new diagnostic criteria groups and that existing diagnostic criteria groups may choose to converge toward this model in their future updates and revisions to incorporate new relevant discoveries. It is also possible to apply this model to rare movement disorders where available diagnostic data may be scarce. In those circumstances, the use of the proposed framework may not lead to specific diagnostic criteria but will immediately inform the field of the strength of any proposed criterion and indicate areas in need of progress. We envision that the presented framework may shape a unified and rigorous, but flexible, approach to diagnostic definition programs. If MDS adopts such an approach, the broad recognition of GRADE methodology across multiple medical specialties globally will also bring additional value and visibility to the society as a global leader in movement disorders. We acknowledge Sarah Wahlstrom Helgren for editorial support in the preparation of this manuscript and the comments and feedback from Francisco Cardoso, Claudia Trenkwalder, Bastiaan R. Bloem, and Louis Chew-Seng Tan. The Rush Program is a Center of Clinical Excellence of the Parkinson's Foundation. T.A.M.: 1. Research project: of the first and 1. Research project: of the first and 1. Research project: and 1. Research project: and 1. Research project: and 1. Research project: and T.A.M.: of for the past has personal for as for and has personal on a for and the International Parkinson and Movement has research support from the of the Research The Parkinson and the of Research of for the past to from and of for the past or membership with of of International Parkinson and Movement and Parkinson's Foundation. of for the past and membership with and and of of International Parkinson and Movement and The for Parkinson's International Parkinson and Movement of The for Parkinson's and of and of for the past or membership with for to Rush Center from of and The for research by from the International Parkinson and Movement by program sponsored by and Rush of for the past or membership with the and from not new data or the research. of methodologies adopted for development of diagnostic criteria of various neurodegenerative of core clinical features in diagnostic criteria of various neurodegenerative The is not responsible for the content or of any information by the than should be to the for the

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.545
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.065
GPT teacher head0.338
Teacher spread0.273 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations7
Published2023
Admission routes2
Has abstractyes

Explore more

Same venueMovement DisordersSame topicParkinson's Disease Mechanisms and TreatmentsFrench-language works237,207