MétaCan
Menu
Back to cohort
Record W2959750635 · doi:10.1038/s41598-019-46649-z

Comparative performances of machine learning methods for classifying Crohn Disease patients using genome-wide genotyping data

2019· article· en· W2959750635 on OpenAlexaff
Alberto Romagnoni, Simon Jégou, Kristel Van Steen, Gilles Wainrib, Jean‐Pierre Hugot, Laurent Peyrin‐Biroulet, Mathias Chamaillard, Jean-Frederick Colombel, Mario Cottone, Mauro D’Amato, R. D’Incà, Jonas Halfvarson, Paul Henderson, Amir Karban, Nicholas A. Kennedy, Mohammed Azam Khan, Marc Lémann, Arie Levine, Dunecan Massey, Mónica Milla, Sok Meng Evelyn Ng, Ioannis Oikonomou, Harald Peeters, Deborah D. Proctor, Jean‐François Rahier, Paul Rutgeerts, Frank Seibold, Laura Stronati, Kirstin M. Taylor, Leif Törkvist, Kullak Ublick, Johan Van Limbergen, A. Van Gossum, Morten H. Vatn, Hu Zhang, Wei Zhang, Jane M. Andrews, Peter A. Bampton, Murray L. Barclay, Timothy H. Florin, Richard B. Gearry, Krupa Krishnaprasad, Ian C. Lawrance, Gillian Mahy, Grant W. Montgomery, Graham Radford-Smith, Rebecca L. Roberts, Lisa A. Simms, Katherine Hanigan, Anthony Croft, Leila Amininijad, Isabelle Cleynen, Olivier Dewit, Denis Franchimont, Michel Georges, Debby Laukens, Emilie Théâtre, Séverine Vermeire, Guy Aumais, Leonard Baidoo, Arthur Barrie, Karen Beck, Edmond-Jean Bernard, David G. Binion, Alain Bitton, Steven R. Brant, Judy H. Cho, Albert Cohen, Kenneth Croitoru, Mark J. Daly, Lisa W. Datta, Colette Deslandres, Richard H. Duerr, Debra Dutridge, John Ferguson, Joann Fultz, Philippe Goyette, Gordon R. Greenberg, Talin Haritunians, Gilles Jobin, Seymour Katz, Raymond Lahaie, Dermot McGovern, Linda D. Nelson, Kaida Ning, Pierre Paré, Miguel Regueiro, John D. Rioux, Elizabeth Ruggiero, L. Philip Schumm, Marc Schwartz, Regan Scott, Yashoda Sharma, Mark S. Silverberg, D. Ross Spears, A. Hillary Steinhart, Joanne M. Stempak, Jason M. Swoger, Constantina Tsagarelis, Wei Zhang, Clarence Zhang, Hongyu Zhao, Jan Aerts, Tariq Ahmad, Hazel Arbury, Anthony Attwood, Adam Auton, Stephen G. Ball, Anthony J. Balmforth, C. Barnes, Jeffrey C. Barrett, Inês Barroso, Anne Barton, Amanda J. Bennett, Sanjeev S. Bhaskar, Katarzyna Błaszczyk, John Bowes, Stephan Brand, Peter S. Braund, Francesca Bredin, Gerome Breen, Ian N Bruce, Jaswinder Bull, Oliver S. Burren, John H. Burton, Jake Byrnes, Sian Caesar, Niall J. Cardin, Chris M. Clee, Alison J. Coffey, John Connell, Donald F. Conrad, Anna F. Dominiczak, Kate Downes, Hazel E. Drummond, Darshna Dudakia, Andrew Dunham, Bernadette Ebbs, Diana Eccles, Sarah Edkins, Cathryn Edwards, Anna Elliot, Paul Emery, David M. Evans, D. Gareth Evans, Anne Farmer, I. Nicol Ferrier, Edward Flynn, Alistair Forbes, Liz Forty, Jayne A. Franklyn, Timothy M. Frayling, Rachel M. Freathy, Eleni Giannoulatou, Polly Gibbs, Paul Gilbert, Katherine Gordon‐Smith, Emma Gray, Elaine Green, Chris Groves, Detelina Grozeva, Rhian Gwilliam, Anita Hall, Naomi Hammond, Matt Hardy, Neelam Hassanali, Husam Hebaishi, Sarah Hines, Anne Hinks, G. A. Hitman, Lynne J. Hocking, Chris Holmes, Eleanor Howard, Philip Howard, Joanna M. M. Howson, Debbie Hughes, Sarah Hunt, John D. Isaacs, Mahim Jain, Derek P. Jewell, Toby Johnson, Jennifer D M Jolley, Ian Jones, Lisa Jones, George Kirov, Cordelia F. Langford, Hana Lango Allen, G.M. Lathrop, James Lee, Kate Lee, Charlie W. Lees, Kevin Lewis, Cecilia M. Lindgren, Meeta Maisuria-Armer, Julian Maller, John Mansfield, Jonathan L. Marchini, Paul Martin, Wendy L. McArdle, Peter McGuffin, Kirsten McLay, Gil McVean, Alexander J. Mentzer, Michael L. Mimmack, Andrew P. Morris, Craig Mowat, Patricia B. Munroe, Simon Myers, William G. Newman, Elaine R. Nimmo, Michael O‘Donovan, Abiodun Onipinla, Nigel Ovington, Michael J. Owen, Kimmo Palin, Aarno Palotie, Kirstie Parnell, Richard D. Pearson, David Pernet, John R. B. Perry, Anne Phillips, Vincent Plagnol, Natalie J. Prescott, Inga Prokopenko, Michael A. Quail, Suzanne Rafelt, Nigel W. Rayner, David M. Reid, Anthony Renwick, Susan M. Ring, Neil M. Robertson, Samuel C. Robson, Ellie Russell, David St Clair, Jennifer G. Sambrook, Jeremy Sanderson, Stephen Sawcer, Helen Schuilenburg, Richard Scott, Sheila Seal, Sue Shaw‐Hawkins, Beverley M. Shields, Matthew J. Simmonds, Debbie J. Smyth, Elilan Somaskantharajah, Katarina Spanova, Sophia Steer, Jonathan Stephens, Helen E. Stevens, Kathy Stirrups, Millicent Stone, David P. Strachan, Zhan Su, Deborah Symmons, John R. Thompson, Wendy Thomson, Martin D. Tobin, Mary E. Travers, Clare Turnbull, Damjan Vukcevic, Louise V. Wain, Mark Walker, Neil Walker, Chris Wallace, Margaret Warren-Perry, Nicholas A. Watkins, John Webster, Michael N. Weedon, Anthony G. Wilson, Matthew Woodburn, B P Wordsworth, Christopher Yau, Allan H. Young, Eleftheria Zeggini, Matthew A. Brown, Paul R. Burton, Mark J. Caulfield, Alastair Compston, Martin Farrall, Stephen Gough, Alistair S. Hall, Andrew T. Hattersley, Adrian V. S. Hill, Christopher G. Mathew, Marcus Pembrey, Jack Satsangi, Michael R. Stratton, Jane Worthington, Matthew E. Hurles, Audrey Duncanson, Willem H. Ouwehand, Miles Parkes, Nazneen Rahman, John A. Todd, Dominic Kwiatkowski, Mark I. McCarthy, Nick Craddock, Panos Deloukas, Peter Donnelly, Jenefer M. Blackwell, Elvira Bramon, Juan P. Casas, Aiden Corvin, Janusz Jankowski, Hugh Markus, Robert Plomin, Anna Rautanen, Richard C. Trembath, Ananth C. Viswanathan, Nicholas Wood, Chris C. A. Spencer, Gavin Band, Céline Bellenguez, Colin Freeman, Garrett Hellenthal, Matti Pirinen, Amy Strange, Hannah Blackburn, Suzannah J. Bumpstead, Serge Dronov, Matthew Gillman, Alagurevathi Jayakumar, Owen T McCann, Jennifer Liddle, Simon Potter, Rathi Ravindrarajah, Michelle Ricketts, Matthew Waller, Paul A. Weston, Sara Widaa, Pamela Whittaker

Bibliographic record

VenueScientific Reports · 2019
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicInflammatory Bowel Disease
Canadian institutionsUniversity of British Columbia HospitalUniversité LavalHôpital Saint-LucCollège de MaisonneuveUniversity of British ColumbiaMontreal Heart InstituteCentre Hospitalier Universitaire Sainte-JustineSt. Michael's HospitalJewish General HospitalUniversity of TorontoRoyal Victoria HospitalMcGill University Health CentreRoyal Victoria Regional Health CentreHospital for Sick ChildrenUniversité de MontréalHôpital Maisonneuve-RosemontMount Sinai HospitalHôtel-Dieu de Montréal
FundersNational Institute of Diabetes and Digestive and Kidney DiseasesAgence Nationale de la RechercheNational Cancer InstituteMedical Research CouncilNational Institute for Health and Care ResearchLaboratoire d'Excellence Inflamex
KeywordsGenome-wide association studyImputation (statistics)Artificial intelligenceInflammatory bowel diseaseMachine learningGenetic associationGenetic architectureComputer scienceMissing dataComputational biologyBiologyDiseaseMedicineQuantitative trait locusGenotypeSingle-nucleotide polymorphismGeneticsInternal medicineGene

Abstract

fetched live from OpenAlex

Crohn Disease (CD) is a complex genetic disorder for which more than 140 genes have been identified using genome wide association studies (GWAS). However, the genetic architecture of the trait remains largely unknown. The recent development of machine learning (ML) approaches incited us to apply them to classify healthy and diseased people according to their genomic information. The Immunochip dataset containing 18,227 CD patients and 34,050 healthy controls enrolled and genotyped by the international Inflammatory Bowel Disease genetic consortium (IIBDGC) has been re-analyzed using a set of ML methods: penalized logistic regression (LR), gradient boosted trees (GBT) and artificial neural networks (NN). The main score used to compare the methods was the Area Under the ROC Curve (AUC) statistics. The impact of quality control (QC), imputing and coding methods on LR results showed that QC methods and imputation of missing genotypes may artificially increase the scores. At the opposite, neither the patient/control ratio nor marker preselection or coding strategies significantly affected the results. LR methods, including Lasso, Ridge and ElasticNet provided similar results with a maximum AUC of 0.80. GBT methods like XGBoost, LightGBM and CatBoost, together with dense NN with one or more hidden layers, provided similar AUC values, suggesting limited epistatic effects in the genetic architecture of the trait. ML methods detected near all the genetic variants previously identified by GWAS among the best predictors plus additional predictors with lower effects. The robustness and complementarity of the different methods are also studied. Compared to LR, non-linear models such as GBT or NN may provide robust complementary approaches to identify and classify genetic markers.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.195
Threshold uncertainty score0.615

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.048
GPT teacher head0.346
Teacher spread0.299 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations106
Published2019
Admission routes1
Has abstractyes

Explore more

Same venueScientific ReportsSame topicInflammatory Bowel DiseaseFrench-language works237,207