MétaCan
Menu
← Back to cohort
Record W7115883470 · doi:10.2196/85069

Machine Learning for Estimating Cardiorespiratory Fitness in Patients With Obesity: Protocol for a Retrospective and Prospective Multicenter Cohort Study

2025· article· en· W7115883470 on OpenAlexvenueno aff

Bibliographic record

VenueJMIR Research Protocols · 2025
Typearticle
Languageen
FieldMedicine
TopicCardiovascular and exercise physiology
Canadian institutionsnot available
Fundersnot available
KeywordsCardiorespiratory fitnessProtocol (science)Relevance (law)Health careCohort studyEstimationProspective cohort studyMEDLINE

Abstract

fetched live from OpenAlex

Background: Cardiorespiratory fitness (CRF) is a key predictor of cardiovascular and other health-related diseases in individuals with obesity. CRF is most accurately assessed through maximal exercise testing with advanced gas-analysis equipment (maximum volume of oxygen [VO2max]); however, this approach is time-consuming, costly, and requires specialized expertise. Therefore, submaximal tests and self-reported physical activity levels have been used to develop predictive algorithms to estimate CRF, yet they often performed poorly in individuals with low CRF levels, such as patients with obesity, because they are predominantly developed using data from healthy populations. Studies using machine learning (ML) models based on VO2max data from patients with obesity appear to be lacking in the literature. ML models based on routinely collected clinical measures may offer a more practical and potentially accurate way to estimate CRF, reducing time, costs, and clinical burden. Objective: The primary aim of this study is to use multicenter, longitudinal, real-world clinical data from a uniquely characterized population with obesity to develop and validate a clinically relevant ML model for estimating CRF and to compare its performance with the gold standard of VO2max testing. Methods: A retrospective data set combining assessments of VO2max tests and clinical parameters from adult patients with severe obesity BMI (≥40.0 kg/m2 or 35.0-39.9 kg/m2 with at least 1 obesity-related comorbidity) from Vestfold Hospital Trust, Muritunet Rehabilitation Institution, and Norwegian School of Sport Sciences will be the foundation for developing the ML model. The clinically relevant ML model for estimating CRF will be presented as a web application, allowing easy access and interaction. The model's estimations will be compared against direct VO2max measurements obtained from medical equipment across institutions as part of a prospective validation. Ethical approval has been obtained for the use of 2 databases in the initial model development; approval for the remaining data and prospective phase is pending. Results: Vestfold Hospital Trust, the Norwegian School of Sport Medicine, and Muritunet Rehabilitation Institution conducted more than 2623 VO₂max tests and collected clinical parameters from 1279 adults with severe obesity during 2013-2025, both before, during, and after lifestyle interventions. The first scientific publication of the clinically relevant ML model is expected to be published in 2026. The results of the overall project are expected to be completed in 2028. The project was awarded salary funding in two rounds from Vestfold Hospital (October and December 2023) and salary funding from the Research Council of Norway (October 2024). In addition, the project received allocated supervision hours from InnoMed Norway (April 2024). Conclusions: This project aims to develop a clinically relevant ML model, which serves as a cost-effective tool for CRF estimation in individuals with obesity, improving accessibility to this important health marker. To our knowledge, this is the first initiative in Norway to estimate CRF in individuals with obesity using ML, based on a unique clinical database. The project carries substantial societal value and holds national and international relevance for health care practice and patient outcomes.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.031
metaresearch head score (Gemma)0.019
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Protocol · Consensus signal: Protocol
Teacher disagreement score0.031
Threshold uncertainty score0.165

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0310.019
Meta-epidemiology (narrow)0.0030.002
Meta-epidemiology (broad)0.0030.004
Bibliometrics0.0020.002
Science and technology studies0.0030.001
Scholarly communication0.0020.001
Open science0.0030.002
Research integrity0.0020.003
Insufficient payload (model declined to judge)0.0130.003

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.060
GPT teacher head0.473
Teacher spread0.413 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreProtocol

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Research Protocols→Same topicCardiovascular and exercise physiology→French-language works237,207→