MétaCan
Menu
Back to cohort
Record W4406777767 · doi:10.1097/cm9.0000000000003440

Leveraging artificial intelligence in the fight against aortic calcification

2025· article· en· W4406777767 on OpenAlexaboutno aff
Jingyue Zhou, Yu Wang, Yang Liu, Lifei Ma, Jinhua Cui, Lanlan Zhang, Xiaoqiang Tang

Bibliographic record

VenueChinese Medical Journal · 2025
Typearticle
Languageen
FieldMedicine
TopicParathyroid Disorders and Treatments
Canadian institutionsnot available
FundersNatural Science Foundation of Sichuan ProvinceSichuan UniversityNational Natural Science Foundation of China
KeywordsMedicineAbdominal aortic aneurysmIntensive care medicineInterpretabilityArtificial intelligenceRadiologyComputer scienceAneurysm

Abstract

fetched live from OpenAlex

To the Editor: Artificial intelligence (AI) is revolutionizing the biomedical field by enabling advanced data analysis, predictive modeling, and personalized medicine, driving breakthroughs in diagnosis, treatment, and drug discovery. In pursuit of this goal, researchers are attempting to develop AI-based algorithms and establish models for use in clinical settings. Key challenges in this pursuit include ensuring the models’ accuracy and consistency and addressing issues such as the interpretability of AI decisions, integration into existing clinical workflows, and ethical considerations like data privacy. Additionally, the AI model lies in the quality and diversity of training data—robust models require diverse and representative datasets to ensure generalizability across different patient populations, reduce dependence on extensive labeled data, and remain resilient to domain shifts, enabling adaptation to new and unseen cases. Nevertheless, this field continues to grow, especially in image-based AI models for diagnosing diseases, such as cardiovascular diseases. Cardiovascular diseases are among the highest risk factors for mortality, and their incidence rates continue to increase.[1] Abdominal aortic calcification (AAC) is exceedingly prevalent in older adults and is closely associated with major adverse cardiovascular events (MACEs). AAC is the ectopic deposition of calcium phosphate minerals in the abdominal aorta that affects the medical or surgical treatment and prognosis of vascular diseases. Epidemiological evidence from the China Dialysis Calcification Study, a prospective cohort study, determined the prevalence of vascular calcification in patients with chronic kidney disease–mineral and bone disorders (n = 1493) to be 77.4% for patients diagnosed with vascular calcification, and 46.8% for patients diagnosed with AAC.[2] Current strategies for quantifying aortic calcification include electron-beam computed tomography (EBCT), computed tomography (CT), and plain radiography. Generally, AAC is assessed using lateral lumbar spine radiographs and lateral spine scans obtained by dual-energy X-ray absorptiometry (DXA) during community healthcare screening. AAC can be visually evaluated by trained imaging specialists using a 24-point system abdominal aortic calcification-24 model (AAC-24) based on the linear length of the calcified aortic wall relative to the height of the lumbar vertebrae. Once AAC is suspected based on plain radiography, CT or EBCT is recommended to verify the diagnosis and to better understand the extent of AAC. This process requires highly trained personnel and proficient specialists for further examination. However, this specialist-based approach is inaccessible to many hospitals due to the limited number of specialists and the relatively high cost of examinations. Thus, the development of upgraded and sophisticated clinical strategies using AI models with sufficient sensitivity and specificity is crucial to addressing conventional challenges and becoming a potential diagnosis approach. Consequently, several studies have attempted to use AI to diagnose cardiovascular diseases. Machine learning (ML) is a typical computational algorithm of AI, and there are two main types of ML models assisting in diagnosing aortic calcification. One is a biomarker-based ML model that detects specific proteins, cytokines, and extracellular vesicle contents, and the other is an image-based AI/ML model that assembles data from CT, X-ray, magnetic resonance imaging, and other imaging modalities. Biomarker-based ML models focus on analyzing biomarkers in the blood samples. By training ML algorithms with large datasets containing biomarker information and clinical parameters, disease progression can be predicted. This approach has proven particularly helpful in stratifying patient risk, allowing clinicians to decide whether further imaging or more aggressive management is warranted. Recent advances in biomarker discovery and multi-omics technologies, such as integrated proteomics and genomics, have improved the reliability of these methods. While accuracy varies across diseases, an ML-based combination of multiple markers has significantly enhanced CVD diagnostics. The other is the image-based ML model. Unlike traditional methods that depend heavily on manual assessment by radiologists, image-based ML models can automatically identify and quantify calcifications. In recent years, convolutional neural networks (CNNs), a type of deep learning model, have been effectively used for detecting and grading aortic calcification from CT and radiographic images. These CNN-based models have shown remarkable accuracy in delimiting the severity of calcifications, often outperforming or matching experienced imaging specialists in consistency and precision.[3] Nevertheless, image-based ML models face limitations in disease screening, particularly for aortic calcification. Arterial imaging is rarely performed on asymptomatic individuals without clinical signs, resulting in a scarcity of imaging data from individuals without arterial disease. Consequently, alternative strategies are needed to more effectively predict or screen patients at risk of aortic calcification. An upgraded model using state-of-the-art ML was recently reported to analyze AAC using lateral spinal images [Figure 1].[4] This new model uses the advanced EfficientNet and generates visual explanations in the form of localization heat maps to identify and rapidly assess the extent of AAC. The authors first trained the established model and tested the performance of the machine learning-abdominal aortic calcification-24 model (ML-AAC-24) scores using 5012 thoracolumbar lateral spine dual-energy X-ray absorptiometry (DXA) images from two manufacturers (Hologic [San Diego, California, USA] and General Electric Company [Cincinnati, Ohio, USA]). They found a relatively high consistency between the ML-AAC-24 and AAC-24 scores provided by trained imaging specialists. The authors then tested the association between the ML-AAC-24 and 15-year cardiovascular disease/all-cause mortality in 1082 women, proving its relatively high application value in clinical settings. Subsequently, the model was validated in a real-world assessment using a registry-based cohort study involving 8565 older adults, primarily women. Cox proportional hazard models were used to identify the relationship between ML-AAC-24 scores and MACE. The Kaplan–Meier event-free survival for MACE and their components (all-cause mortality, acute myocardial infarction, and ischemic cerebrovascular events) showed an increasing divergence from the observation onset to the end of follow-up. Thus, the ML-AAC-24 assay is reliable. Based on these investigations, an EfficientNet model, incorporating advanced ML algorithms, was established to reliably identify and assess AAC progression using the AAC-24 system.Figure 1: Research and application flowchart for leveraging AI in the fight against aortic calcification. The ML-AAC-24 was established with enormous training in cohorts and was carefully verified. Its high consistency and accuracy among imaging specialists may promote its application in clinical settings. The figure was created using BioRender (www.BioRender.com). AAC-24: Abdominal aortic calcification-24 model; AI: Artificial intelligence; APP: Application program; MACE: Major adverse cardiovascular events; ML-AAC-24: Machine learning-abdominal aortic calcification-24 model.CT and EBCT are the most common and efficient methods for detecting AAC in clinical practice; however, they are time-consuming and expensive. The ML-AAC-24 model requires less time to estimate AAC-24 scores and exhibits relatively high agreement and accuracy among imaging specialists. Sharif et al[4] reported that the intraclass correlation coefficient between the ML-AAC-24 and the AAC-24 scores provided by imaging specialists for all test sets was 0.84, with a Pearson correlation coefficient of 0.86, indicating excellent agreement during clinical use. Furthermore, the average classification accuracy was 80% for the three AAC groups (low, moderate, and high), which is expected to increase in future interactions of the model. Subsequently, the ML-AAC-24 was used to evaluate the association between the ML-AAC-24 and the long-term incidence of falls and fractures in 1023 women in the Perth Longitudinal Study of Aging Women.[5] Notably, the ML-AAC-24 results did not differ from the manual estimation results, demonstrating that EfficientNet produced significantly better results than the canonical Bayesian model. Moreover, another ML-based algorithm using an advanced U-Net model was superior to EfficientNet, with a Pearson’s correlation coefficient of 0.97.[6] Therefore, these ML-based models may be more affordable and less labor-intensive alternatives to evaluate the risk of cardiovascular disease and to rapidly assess the extent of AAC from easy-to-obtain lateral spine images [Supplementary Table 1, https://links.lww.com/CM9/C269]. Despite these promising advances, some questions remain and require further investigation. (1) Risk factors: AAC has many risk factors that may affect the agreement and accuracy of the ML-AAC-24 model. For instance, compared with women, men are more likely to develop vascular calcification;[7] however, the Perth longitudinal study was conducted among women, and the Manitoba registry mainly comprised of women. Therefore, future studies should focus on the agreement and accuracy of the ML-AAC-24 in male cohorts, such as the Osteoporotic Fractures in Men Study (MrOS), a 7-year prospective observational study investigating the risk factors for osteoporosis and fracture. Importantly, this study measured spinal bone mineral density, meeting the requirements for the ML-AAC-24. Moreover, multicenter trials involving multiple ethnic groups and patients with complex genetic backgrounds are warranted. Other traditional risk factors, such as smoking, hyperlipidemia, hypertension, and diabetes mellitus, could also affect the predictive efficiency of AAC. Thus, stratification analyses should be conducted in real-world assessments to improve predictive efficiency. (2) Algorithm optimizations: Although the accuracy of the ML-AAC-24 was relatively high, some individuals received inaccurate assessments due to specific interferences. In some cases, interference with the observation area occurred, including extensive bowel gas overlay (extensively), hyperlordosis-induced rib malposition (adjacent to the L1 region), and iliac artery calcification (adjacent to the L4 region). In addition, vertebral fractures, scoliosis, osteophytes, rotation, and less extensive AAC may interfere with estimation. These misdiagnoses might be improved by deep adaptive graph-dependent anatomical landmark detection which exhibits excellent performance in automated bone mineral density predictions and increases accuracy. In addition, the ML-AAC-24 tended to underestimate the AAC-24 scores. Therefore, optimizing the models, improving image preprocessing, and adjusting hyperparameters could address this issue. (3) Clinical application: Higher accuracy and broader applicability are expected after equipment updates and algorithm optimization. In theory, individuals or doctors can upload plain radiographic results to the associated software or Applications (APPs). The ML-AAC-24 could rapidly report the degree of AAC risk and offer personalized prevention and treatment suggestions [Figure 1]. However, this vision requires optimization of ML-AAC-24 scores in a real-world study. (4) Development strategy: Considering the black-box nature of the ML algorithm, it is difficult for clinicians to understand how the model makes decisions and subsequently trust the results provided by the ML-AAC-24. In this case, changing the goal of the ML-AAC-24 from “close to specialist” to “best assistant to specialists” may be a promising development strategy. This can be achieved by transforming the visual image presented to clinicians from localization heat maps (illustrating the model’s qualitative performance) to a schematic diagram of the thoracolumbar lateral spine images. This mode is more visual and conducive for imaging specialists to assess. In addition, arterial calcification is a systemic pathological process, and biomarker-based ML models are limited to AAC evaluation. Although some miRNAs have been reported as potential biomarkers of vascular calcification, accurate biomarkers for site-specific arterial calcification are still under investigation. However, a combined (i.e., biomarker-image) ML model can improve this situation. Further studies should integrate the biomarker-based model with the image-based ML model to detect global arterial calcification levels and simultaneously evaluate the risk of AAC in community health assessments [Supplementary Figure 1, https://links.lww.com/CM9/C269]. With the indication of biomarker levels, the combined ML models detected AAC on lateral spine scans obtained from DXA using ML-AAC-24. Further examination is recommended when the predicted risk of the models is high. Thus, AAC could be better evaluated simultaneously with global arterial calcification detection during community health assessments. Using these approaches, the ML-AAC-24 could be “the best assistant for specialists”, especially considering the large workload associated with community health assessments, as it better fits current clinical settings and lessens the workload of imaging specialists. Funding This work was supported by the Sichuan Natural Science Foundation Outstanding Youth Science Foundation (No. 2024NSFJQ0053), the National Natural Science Foundation of China (No. 82370235), the Tianfu Qingcheng Plan (No. 1711), and the K-funding of West China Second University Hospital Sichuan University (grant No. KZ197). Conflicts of interest None.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.733
Threshold uncertainty score0.251

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.025
GPT teacher head0.356
Teacher spread0.331 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueChinese Medical JournalSame topicParathyroid Disorders and TreatmentsFrench-language works237,207