MétaCan
Menu
Back to cohort
Record W4380151099 · doi:10.1101/2023.06.06.23290887

GestaltMatcher Database - A global reference for facial phenotypic variability in rare human diseases

2023· preprint· en· W4380151099 on OpenAlexaff
Hellen Lesmann, Alexander Hustinx, Shahida Moosa, Hannah Klinkhammer, Elaine Marchi, Pilar Caro, Ibrahim M. Abdelrazek, Jean Tori Pantel, Merle ten Hagen, Meow‐Keong Thong, Rifhan Azwani Mazlan, Sok Kun Tae, Tom Kamphans, Wolfgang Meiswinkel, Jingmei Li, Behnam Javanmardi, Alexej Knaus, Annette Uwineza, Cordula Knopp, Tinatin Tkemaladze, Miriam Elbracht, Larissa Mattern, Rami Abou Jamra, Clara Velmans, Vincent Strehlow, Maureen Jacob, Angela Peron, Cristina Dias, Beatriz Nunes, Thainá Vilella, Isabel Furquim Pinheiro, Chong Ae Kim, Maria Isabel Melaragno, Hannah Weiland, Sophia Kaptain, Karolina Chwiałkowska, Mirosław Kwaśniewski, Ramy Saad, Sarah Wiethoff, Himanshu Goel, Clara Sze-Man Tang, Anna Hau, Tahsin Stefan Barakat, Przemysław Panek, Amira Nabil, Julia Suh, Frederik Braun, Israel Gomy, Luisa Averdunk, Ekanem N. Ekure, Gaber Bergant, Borut Peterlin, Claudio Graziano, Nagwa E. A. Gaboon, Moisés Ó. Fiesco-Roa, Alessandro Spinelli, Nina‐Maria Wilpert, Prasit Phowthongkum, Nergis Güzel, Tobias B. Haack, Rana Bitar, Andreas Tzschach, Agustí Rodríguez‐Palmero, Theresa Brunet, Sabine Rudnik‐Schöneborn, Silvina Contreras‐Capetillo, Ava Oberlack, Carole Samango‐Sprouse, Teresa Sadeghin, Margaret Olaya, Konrad Platzer, Artem Borovikov, Franziska Schnabel, Lara Heuft, Vera Herrmann, Renske Oegema, Nour Elkhateeb, Sheetal Kumar, Katalin Komlósi, Khoushoua Mohamed, Silvia Kalantari, Fabio Sirchia, Antonio F. Martinez-Monseny, Matthias Höller, Louiza Toutouna, Amal Mohamed, Amaia Lasa‐Aranzasti, John A. Sayer, Nadja Ehmke, Magdalena Danyel, Henrike L. Sczakiel, Sarina Schwartzmann, Felix Boschann, Max Zhao, R. Adam, Lara Einicke, Denise Horn, Kee Seang Chew, KAM Choy Chen, Miray Karakoyun, Ben Pode‐Shakked, Aviva Eliyahu, Rachel Rock, Teresa Carrion, Odelia Chorin, Yuri A. Zárate, Marcelo Martinez Conti, Mert Karakaya, Moon Ley Tung, Bharatendu Chandra, Arjan Bouman, Aimé Lumaka, Naveed Wasif, Marwan Shinawi, Patrick R. Blackburn, Tianyun Wang, Tim Niehues, Axel Schmidt, Regina Roth, Dagmar Wieczorek, Ping Hu, Rebekah L. Waikel, Suzanna E. Ledgister Hanchard, Gehad Elmakkawy, Sylvia Safwat, Frédéric Ebstein, Elke Krüger, Sébastien Küry, Stéphane Bezieau, Annabelle Arlt, Eric Olinger, Felix Marbach, Dong Li, Lucie Dupuis, Roberto Mendoza‐Londono, Sofia Douzgou, Denisa Weis, Brian Hon‐Yin Chung, Christopher C.Y. Mak, Hülya Kayserili, Nursel Elçioğlu, Ayça Aykut, Peli Özlem Şimşek-Kiper, Nina Bögershausen, Bernd Wollnik, Heidi Beate Bentzen, Ingo Kurth, Christian Netzer, Aleksandra Jezela‐Stanek, Koenraad Devriendt, Karen W. Gripp, Martin Mücke, Alain Verloès, Christian P. Schaaf, Christoffer Nellåker, Benjamin D. Solomon, Markus M. Nöthen, Ebtesam Abdalla, Gholson J. Lyon, Peter Krawitz, Tzung‐Chien Hsieh

Bibliographic record

VenuemedRxiv · 2023
Typepreprint
Languageen
FieldComputer Science
TopicAI in cancer detection
Canadian institutionsHospital for Sick ChildrenHealth Research Foundation
FundersNational Human Genome Research InstituteNational Institutes of Health
KeywordsBenchmarkingInteroperabilityMedicineComputer scienceDatabaseWorld Wide Web

Abstract

fetched live from OpenAlex

The most important factor that complicates the work of dysmorphologists is the significant phenotypic variability of the human face. Next-Generation Phenotyping (NGP) tools that assist clinicians with recognizing characteristic syndromic patterns are particularly challenged when confronted with patients from populations different from their training data. To that end, we systematically analyzed the impact of genetic ancestry on facial dysmorphism. For that purpose, we established the GestaltMatcher Database (GMDB) as a reference dataset for medical images of patients with rare genetic disorders from around the world. We collected 10,980 frontal facial images - more than a quarter previously unpublished - from 8,346 patients, representing 581 rare disorders. Although the predominant ancestry is still European (67%), data from underrepresented populations have been increased considerably via global collaborations (19% Asian and 7% African). This includes previously unpublished reports for more than 40% of the African patients. The NGP analysis on this diverse dataset revealed characteristic performance differences depending on the composition of training and test sets corresponding to genetic relatedness. For clinical use of NGP, incorporating non-European patients resulted in a profound enhancement of GestaltMatcher performance. The top-5 accuracy rate increased by +11.29%. Importantly, this improvement in delineating the correct disorder from a facial portrait was achieved without decreasing the performance on European patients. By design, GMDB complies with the FAIR principles by rendering the curated medical data findable, accessible, interoperable, and reusable. This means GMDB can also serve as data for training and benchmarking. In summary, our study on facial dysmorphism on a global sample revealed a considerable cross ancestral phenotypic variability confounding NGP that should be counteracted by international efforts for increasing data diversity. GMDB will serve as a vital reference database for clinicians and a transparent training set for advancing NGP technology.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.781
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0020.002
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.078
GPT teacher head0.337
Teacher spread0.259 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations24
Published2023
Admission routes1
Has abstractyes

Explore more

Same venuemedRxivSame topicAI in cancer detectionFrench-language works237,207