MétaCan
Menu
Back to cohort
Record W4386776358 · doi:10.1016/j.eclinm.2023.102210

Characterization of long COVID temporal sub-phenotypes by distributed representation learning from electronic health record data: a cohort study

2023· article· en· W4386776358 on OpenAlexaff
Arianna Dagliati, Zachary H. Strasser, Zahra Shakeri Hossein Abad, Jeffrey G. Klann, Kavishwar B. Wagholikar, Rebecca Mesa, Shyam Visweswaran, Michele Morris, Yuan Luo, Darren W. Henderson, Malarkodi Jebathilagam Samayamuthu, Bryce W. Q. Tan, Guillame Verdy, Gilbert S. Omenn, Zongqi Xia, Riccardo Bellazzi, James R. Aaron, Giuseppe Agapito, Adem Albayrak, Giuseppe Albi, M Alessiani, Anna Alloni, Danilo F. Amendola, François Angoulvant, Li L.L.J. Anthony, Bruce J. Aronow, Fatima Ashraf, Andrew M. Atz, Paul Avillach, Paula S. Azevedo, James Balshi, Brett K. Beaulieu‐Jones, Douglas S. Bell, Antonio Bellasi, Vincent Benoît, Michele Beraghi, José Luis Bernal-Sobrino, Mélodie Bernaux, Romain Bey, Surbhi Bhatnagar, Alvar Blanco-Martínez, Clara-Lea Bonzel, John Booth, Silvano Bosari, Florence T. Bourgeois, Robert L. Bradford, Gabriel A. Brat, Stéphane Breant, Nicholas W. Brown, Raffaele Bruno, William Bryant, Mauro Bucalo, Emily M. Bucholz, Anita Burgun, Tianxi Cai, Mario Cannataro, Aldo Carmona, Charlotte Caucheteux, Julien Champ, Jin Chen, Krista Y. Chen, Luca Chiovato, Lorenzo Chiudinelli, Kelly Cho, James J. Cimino, Tiago K. Colicchio, Sylvie Cormont, Sébastien Cossin, Jean B. Craig, Juan Luis Cruz-Bermúdez, Jaime Cruz‐Rojo, Mohamad Daniar, Christel Daniel, Priyam Das, Batsal Devkota, Audrey Dionne, Rui Duan, Julien Dubiel, Scott L. DuVall, Loïc Estève, Hossein Estiri, Shirley Fan, Robert W Follett, Thomas Ganslandt, Noelia García Barrio, Lana X. Garmire, Nils Gehlenborg, Emily Getzen, Alon Geva, Tobias Gradinger, Alexandre Gramfort, Romain Griffier, Nicolas Griffon, Olivier Grisel, Alba Gutiérrez‐Sacristán, Larry Han, David A. Hanauer, Christian Haverkamp, Derek Hazard, Bing He, Martin Hilka, Yuk‐Lam Ho, John H. Holmes, Chuan Hong, Kenneth M. Huling, Meghan R. Hutch, Richard Issitt, Anne‐Sophie Jannot, Vianney Jouhet, Ramakanth Kavuluru, Mark S. Keller, Chris J. Kennedy, Daniel Key, Katie Kirchoff, Isaac S. Kohane, Ian D. Krantz, Detlef Kraska, Ashok Krishnamurthy, Sehi L'Yi, Trang T. Le, Judith Leblanc, Guillaume Lemaître, Leslie Lenert, Damien Leprovost, Molei Liu, Ne Hooi Will Loh, Qi Long, Sara Lozano‐Zahonero, Kristine E. Lynch, Sadiqa Mahmood, Sarah E. Maidlow, Adeline Makoudjou, Alberto Malovini, Kenneth D. Mandl, Chengsheng Mao, Anupama Maram, Patricia Martel, Marcelo Roberto Martins, Jayson S. Marwaha, Aaron J. Masino, Maria Mazzitelli, Arthur Mensch, Marianna Milano, Marcos Ferreira Minicucci, Bertrand Moal, Taha Mohseni Ahooyi, Jason H. Moore, Cinta Moraleda, Jeffrey S. Morris, Karyn Moshal, Sajad Mousavi, Danielle L. Mowery, Douglas A. Murad, Shawn N. Murphy, Thomas P. Naughton, Carlos Tadeu Breda Neto, Antoine Neuraz, Jane W. Newburger, Kee Yuan Ngiam, Wanjikũ Njoroge, James B. Norman, Jihad S. Obeid, Marina Politi Okoshi, Karen L. Olson, Nina Orlova, Brian D. Ostasiewski, Nathan Palmer, Nicolás Paris, Lav P. Patel, Miguel Pedrera‐Jiménez, Emily Pfaff, Ashley C. Pfaff, Danielle Pillion, Sara Pizzimenti, Hans U. Prokosch, Robson A. Prudente, Andrea Prunotto, Víctor Quirós González, Rachel Ramoni, Maryna Raskin, Siegbert Rieg, Gustavo Roig-Domínguez, Pablo Rojo, Paula Rubio-Mayo, Paolo Sacchi, Carlos Sáez, Elisa Salamanca, L. Nelson Sanchez‐Pinto, Arnaud Sandrin, Nandhini Santhanam, Janaina C.C. Santos, Fernando J Sanz Vidorreta, Emily Schriver, Petra Schubert, Juergen Schuettler, Luigia Scudeller, Neil J. Sebire, Pablo Serrano Balazote, Patricia Serre, Arnaud Serret-Larmande, Mohsin Shah, Domenick Silvio, Piotr Sliz, Jiyeon Son, Charles Sonday, Andrew M. South, Anastassia Spiridou, Amelia L.M. Tan, Byorn W.L. Tan, Suzana Érico Tanni, Deanne M. Taylor, Ana I. Terriza-Torres, Valentina Tibollo, Patric Tippmann, Emma M. S. Toh, Carlo Torti, Enrico Maria Trecarichi, Yi‐Ju Tseng, Andrew K. Vallejos, Gaël Varoquaux, Margaret E. Vella, Guillaume Verdy, Jill-Jênn Vie, Shyam Visweswaran, Michele Vitacca, Lemuel R. Waitman, Xuan Wang, Demián Wassermann, Griffin M. Weber, Martin Wolkewitz, Scott Wong, Xin Xiong, Ye Ye, Nadir Yehya, William Yuan, Alberto Zambelli, Harrison G. Zhang, Daniela Zo ̈ller, Valentina Zuccaro, Chiara Zucco

Bibliographic record

VenueEClinicalMedicine · 2023
Typearticle
Languageen
FieldMedicine
TopicLong-Term Effects of COVID-19
Canadian institutionsPublic Health OntarioUniversity of Toronto
FundersNational Center for Advancing Translational SciencesNational Institute of Environmental Health SciencesNational Heart, Lung, and Blood InstituteBritish Heart FoundationNational Medical Research CouncilNational Institute of Neurological Disorders and StrokeNational Institute of Allergy and Infectious DiseasesMedical Research CouncilNational Institutes of HealthEuropean CommissionNational Institute on AgingHorizon 2020 Framework Programme
KeywordsMedicineCoronavirus disease 2019 (COVID-19)Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2)Electronic health record2019-20 coronavirus outbreakCohortPhenotypeRepresentation (politics)Cohort studyHealth recordsPandemicVirologyPathologyGeneticsOutbreakInfectious disease (medical specialty)Health careDisease

Abstract

fetched live from OpenAlex

Background: has been challenging due to the multitude of sub-phenotypes, temporal attributes, and definitions. Scalable characterization of PASC sub-phenotypes can enhance screening capacities, disease management, and treatment planning. Methods: We conducted a retrospective multi-centre observational cohort study, leveraging longitudinal electronic health record (EHR) data of 30,422 patients from three healthcare systems in the Consortium for the Clinical Characterization of COVID-19 by EHR (4CE). From the total cohort, we applied a deductive approach on 12,424 individuals with follow-up data and developed a distributed representation learning process for providing augmented definitions for PASC sub-phenotypes. Findings: Our framework characterized seven PASC sub-phenotypes. We estimated that on average 15.7% of the hospitalized COVID-19 patients were likely to suffer from at least one PASC symptom and almost 5.98%, on average, had multiple symptoms. Joint pain and dyspnea had the highest prevalence, with an average prevalence of 5.45% and 4.53%, respectively. Interpretation: We provided a scalable framework to every participating healthcare system for estimating PASC sub-phenotypes prevalence and temporal attributes, thus developing a unified model that characterizes augmented sub-phenotypes across the different systems. Funding: Authors are supported by National Institute of Allergy and Infectious Diseases, National Institute on Aging, National Center for Advancing Translational Sciences, National Medical Research Council, National Institute of Neurological Disorders and Stroke, European Union, National Institutes of Health, National Center for Advancing Translational Sciences.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.003
metaresearch head score (Gemma)0.008
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.055
Threshold uncertainty score0.990

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0030.008
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.048
GPT teacher head0.396
Teacher spread0.349 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations24
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueEClinicalMedicineSame topicLong-Term Effects of COVID-19French-language works237,207