Prediction Robustness and Data Redundancy in Machine Learning for Materials Science
Notice bibliographique
Résumé
Prediction Robustness and Data Redundancy in Machine Learning for Materials ScienceKangming Li a, Daniel Persaud a, Kamal Choudhary b, Brian DeCost b, Michael Greenwood c, Jason Hattrick-Simpers aa Department of Materials Science and Engineering, University of Toronto, Canadab Material Measurement Laboratory, National Institute of Standards and Technology, USAc Canmet MATERIALS, Natural Resources Canada, CanadaMaterials for Sustainable Development Conference (MATSUS)Proceedings of MATSUS Spring 2024 Conference (MATSUS24)#AI - Automation and Nanomaterials (machine learning, artificial intelligence, robotics, accelerated discovery)Barcelona, Spain, 2024 March 4th - 8thOrganizers: Ivan Infante and Oleksandr VoznyyInvited Speaker, Kangming Li, presentation 151DOI: https://doi.org/10.29363/nanoge.matsus.2024.151Publication date: 18th December 2023The rapid growth of big data in materials science has led to significant advancements in materials property prediction by machine learning (ML) models. However, big data does not necessarily lead to robust prediction performance of ML models. In addition, the issue of information redundancy in materials data has been largely overlook. This talk intends to present an examination of these two correlated challenges related to materials data: prediction robustness and data redundancy. First, we will discuss the challenges in ensuring the prediction robustness of ML models, by showcasing the severe performance degradation when the models are trained on the Materials Project 2018 dataset and tested on the Materials Project 2021 dataset. We will demonstrate the impact of distribution shifts and use tools such as UMAP and query-by-committee to foresee performance degradation and to improve prediction accuracy. Next, we will delve into the issue of data redundancy across large materials datasets, revealing that up to 95% of materials data can be safely removed with little impact on the model performance. We will highlight the application of uncertainty-based active learning algorithms to create smaller but informative datasets, leading to more efficient data acquisition and ML training. By examining these challenges, this talk aims to provide insights into building more efficient and robust materials databases and ML models for accurate and reliable predictions in materials science. References:[1] Li, K., DeCost, B., Choudhary, K. et al. A critical examination of robustness and generalizability of machine learning prediction of materials properties. npj Comput Mater 9, 55 (2023).[2] Li, K., Persaud, D., Choudhary, K. et al. Exploiting redundancy in large materials datasets for efficient machine learning with less data. Nat Commun 14, 7283 (2023).Acknowledgements:The computations were made on the resources provided by the Calcul Quebec, Westgrid, and Compute Ontario consortia in the Digital Research Alliance of Canada (alliancecan.ca), and the Acceleration Consortium (acceleration.utoronto.ca) at the University of Toronto. We acknowledge funding provided by Natural Resources Canada's Office of Energy Research and Development (OERD). © FUNDACIO DE LA COMUNITAT VALENCIANA SCITOnanoGe is a prestigious brand of successful science conferences that are developed along the year in different areas of the world since 2009. Our worldwide conferences cover cutting-edge materials topics like perovskite solar cells, photovoltaics, optoelectronics, solar fuel conversion, surface science, catalysis and two-dimensional materials, among many others.nanoGe Fall MeetingnanoGe Fall Meeting (NFM) is a multiple symposia conference celebrated yearly and focused on a broad set of topics of advanced materials preparation, their fundamental properties, and their applications, in fields such as renewable energy, photovoltaics, lighting, semiconductor quantum dots, 2-D materials synthesis, charge carriers dynamics, microscopy and spectroscopy semiconductors fundamentals, etc.nanoGe Spring MeetingThis conference is a unique series of symposia focused on advanced materials preparation and fundamental properties and their applications, in fields such as renewable energy (photovoltaics, batteries), lighting, semiconductor quantum dots, 2-D materials synthesis and semiconductors fundamentals, bioimaging, etc.International Conference on Hybrid and Organic PhotovoltaicsInternational Conference on Hybrid and Organic Photovoltaics (HOPV) is celebrated yearly in May. The main topics are the development, function and modeling of materials and devices for hybrid and organic solar cells. The field is now dominated by perovskite solar cells but also other hybrid technologies, as organic solar cells, quantum dot solar cells, and dye-sensitized solar cells and their integration into devices for photoelectrochemical solar fuel production.Asia-Pacific International Conference on Perovskite, Organic Photovoltaics and OptoelectronicsThe main topics of the Asia-Pacific International Conference on Perovskite, Organic Photovoltaics and Optoelectronics (IPEROP) are discussed every year in Asia-Pacific for gathering the recent advances in the fields of material preparation, modeling and fabrication of perovskite and hybrid and organic materials. Photovoltaic devices are analyzed from fundamental physics and materials properties to a broad set of applications. The conference also covers the developments of perovskite optoelectronics, including light-emitting diodes, lasers, optical devices, nanophotonics, nonlinear optical properties, colloidal nanostructures, photophysics and light-matter coupling.International Conference on Perovskite Thin Film Photovoltaics Perovskite Photonics and OptoelectronicsThe International Conference on Perovskite Thin Film Photovoltaics Perovskite Photonics and Optoelectronics (NIPHO) is the best place to hear the latest developments in perovskite solar cells as well as on recent advances in the fields of perovskite light-emitting diodes, lasers, optical devices, nanophotonics, nonlinear optical properties, colloidal nanostructures, photophysics and light-matter coupling.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».