MétaCan
Menu
Back to cohort
Record W4380151099 · doi:10.1101/2023.06.06.23290887

GestaltMatcher Database - A global reference for facial phenotypic variability in rare human diseases

2023· preprint· en· W4380151099 on OpenAlexaff
Hellen Lesmann, Alexander Hustinx, Shahida Moosa, Hannah Klinkhammer, Elaine Marchi, Pilar Caro, Ibrahim M. Abdelrazek, Jean Tori Pantel, Merle ten Hagen, Meow‐Keong Thong, Rifhan Azwani Mazlan, Sok Kun Tae, Tom Kamphans, Wolfgang Meiswinkel, Jingmei Li, Behnam Javanmardi, Alexej Knaus, Annette Uwineza, Cordula Knopp, Tinatin Tkemaladze, Miriam Elbracht, Larissa Mattern, Rami Abou Jamra, Clara Velmans, Vincent Strehlow, Maureen Jacob, Angela Peron, Cristina Dias, Beatriz Nunes, Thainá Vilella, Isabel Furquim Pinheiro, Chong Ae Kim, Maria Isabel Melaragno, Hannah Weiland, Sophia Kaptain, Karolina Chwiałkowska, Mirosław Kwaśniewski, Ramy Saad, Sarah Wiethoff, Himanshu Goel, Clara Sze-Man Tang, Anna Hau, Tahsin Stefan Barakat, Przemysław Panek, Amira Nabil, Julia Suh, Frederik Braun, Israel Gomy, Luisa Averdunk, Ekanem N. Ekure, Gaber Bergant, Borut Peterlin, Claudio Graziano, Nagwa E. A. Gaboon, Moisés Ó. Fiesco-Roa, Alessandro Spinelli, Nina‐Maria Wilpert, Prasit Phowthongkum, Nergis Güzel, Tobias B. Haack, Rana Bitar, Andreas Tzschach, Agustí Rodríguez‐Palmero, Theresa Brunet, Sabine Rudnik‐Schöneborn, Silvina Contreras‐Capetillo, Ava Oberlack, Carole Samango‐Sprouse, Teresa Sadeghin, Margaret Olaya, Konrad Platzer, Artem Borovikov, Franziska Schnabel, Lara Heuft, Vera Herrmann, Renske Oegema, Nour Elkhateeb, Sheetal Kumar, Katalin Komlósi, Khoushoua Mohamed, Silvia Kalantari, Fabio Sirchia, Antonio F. Martinez-Monseny, Matthias Höller, Louiza Toutouna, Amal Mohamed, Amaia Lasa‐Aranzasti, John A. Sayer, Nadja Ehmke, Magdalena Danyel, Henrike L. Sczakiel, Sarina Schwartzmann, Felix Boschann, Max Zhao, R. Adam, Lara Einicke, Denise Horn, Kee Seang Chew, KAM Choy Chen, Miray Karakoyun, Ben Pode‐Shakked, Aviva Eliyahu, Rachel Rock, Teresa Carrion, Odelia Chorin, Yuri A. Zárate, Marcelo Martinez Conti, Mert Karakaya, Moon Ley Tung, Bharatendu Chandra, Arjan Bouman, Aimé Lumaka, Naveed Wasif, Marwan Shinawi, Patrick R. Blackburn, Tianyun Wang, Tim Niehues, Axel Schmidt, Regina Roth, Dagmar Wieczorek, Ping Hu, Rebekah L. Waikel, Suzanna E. Ledgister Hanchard, Gehad Elmakkawy, Sylvia Safwat, Frédéric Ebstein, Elke Krüger, Sébastien Küry, Stéphane Bezieau, Annabelle Arlt, Eric Olinger, Felix Marbach, Dong Li, Lucie Dupuis, Roberto Mendoza‐Londono, Sofia Douzgou, Denisa Weis, Brian Hon‐Yin Chung, Christopher C.Y. Mak, Hülya Kayserili, Nursel Elçioğlu, Ayça Aykut, Peli Özlem Şimşek-Kiper, Nina Bögershausen, Bernd Wollnik, Heidi Beate Bentzen, Ingo Kurth, Christian Netzer, Aleksandra Jezela‐Stanek, Koenraad Devriendt, Karen W. Gripp, Martin Mücke, Alain Verloès, Christian P. Schaaf, Christoffer Nellåker, Benjamin D. Solomon, Markus M. Nöthen, Ebtesam Abdalla, Gholson J. Lyon, Peter Krawitz, Tzung‐Chien Hsieh

Bibliographic record

VenuemedRxiv · 2023
Typepreprint
Languageen
FieldComputer Science
TopicAI in cancer detection
Canadian institutionsHospital for Sick ChildrenHealth Research Foundation
FundersNational Human Genome Research InstituteNational Institutes of Health
KeywordsBenchmarkingInteroperabilityMedicineComputer scienceDatabaseWorld Wide Web

Abstract

fetched live from OpenAlex

The most important factor that complicates the work of dysmorphologists is the significant phenotypic variability of the human face. Next-Generation Phenotyping (NGP) tools that assist clinicians with recognizing characteristic syndromic patterns are particularly challenged when confronted with patients from populations different from their training data. To that end, we systematically analyzed the impact of genetic ancestry on facial dysmorphism. For that purpose, we established the GestaltMatcher Database (GMDB) as a reference dataset for medical images of patients with rare genetic disorders from around the world. We collected 10,980 frontal facial images - more than a quarter previously unpublished - from 8,346 patients, representing 581 rare disorders. Although the predominant ancestry is still European (67%), data from underrepresented populations have been increased considerably via global collaborations (19% Asian and 7% African). This includes previously unpublished reports for more than 40% of the African patients. The NGP analysis on this diverse dataset revealed characteristic performance differences depending on the composition of training and test sets corresponding to genetic relatedness. For clinical use of NGP, incorporating non-European patients resulted in a profound enhancement of GestaltMatcher performance. The top-5 accuracy rate increased by +11.29%. Importantly, this improvement in delineating the correct disorder from a facial portrait was achieved without decreasing the performance on European patients. By design, GMDB complies with the FAIR principles by rendering the curated medical data findable, accessible, interoperable, and reusable. This means GMDB can also serve as data for training and benchmarking. In summary, our study on facial dysmorphism on a global sample revealed a considerable cross ancestral phenotypic variability confounding NGP that should be counteracted by international efforts for increasing data diversity. GMDB will serve as a vital reference database for clinicians and a transparent training set for advancing NGP technology.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.008
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.014
Threshold uncertainty score0.047

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.008
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0060.003
Science and technology studies0.0010.000
Scholarly communication0.0020.001
Open science0.0020.004
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0140.013

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.078
GPT teacher head0.337
Teacher spread0.259 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations24
Published2023
Admission routes1
Has abstractyes

Explore more

Same venuemedRxivSame topicAI in cancer detectionFrench-language works237,207