MétaCan
Menu
Back to cohort
Record W2771697124 · doi:10.1021/acs.accounts.7b00490

Discovery of Intermetallic Compounds from Traditional to Machine-Learning Approaches

2017· article· en· W2771697124 on OpenAlexafffund
Anton O. Oliynyk, Arthur Mar

Bibliographic record

VenueAccounts of Chemical Research · 2017
Typearticle
Languageen
FieldMaterials Science
TopicMachine Learning in Materials Science
Canadian institutionsUniversity of Alberta
FundersNatural Sciences and Engineering Research Council of Canada
KeywordsIntermetallicComputer scienceMaterials scienceMetallurgy

Abstract

fetched live from OpenAlex

Conspectus Intermetallic compounds are bestowed by diverse compositions, complex structures, and useful properties for many materials applications. How metallic elements react to form these compounds and what structures they adopt remain challenging questions that defy predictability. Traditional approaches offer some rational strategies to prepare specific classes of intermetallics, such as targeting members within a modular homologous series, manipulating building blocks to assemble new structures, and filling interstitial sites to create stuffed variants. Because these strategies rely on precedent, they cannot foresee surprising results, by definition. Exploratory synthesis, whether through systematic phase diagram investigations or serendipity, is still essential for expanding our knowledge base. Eventually, the relationships may become too complex for the pattern recognition skills to be reliably or practically performed by humans. Complementing these traditional approaches, new machine-learning approaches may be a viable alternative for materials discovery, not only among intermetallics but also more generally to other chemical compounds. In this Account, we survey our own efforts to discover new intermetallic compounds, encompassing gallides, germanides, phosphides, arsenides, and others. We apply various machine-learning methods (such as support vector machine and random forest algorithms) to confront two significant questions in solid state chemistry. First, what crystal structures are adopted by a compound given an arbitrary composition? Initial efforts have focused on binary equiatomic phases AB, ternary equiatomic phases ABC, and full Heusler phases AB 2 C. Our analysis emphasizes the use of real experimental data and places special value on confirming predictions through experiment. Chemical descriptors are carefully chosen through a rigorous procedure called cluster resolution feature selection. Predictions for crystal structures are quantified by evaluating probabilities. Major results include the discovery of RhCd, the first new binary AB compound to be found in over 15 years, with a CsCl-type structure; the connection between “ambiguous” prediction probabilities and the phenomenon of polymorphism, as illustrated in the case of TiFeP (with TiNiSi- and ZrNiAl-type structures); and the preparation of new predicted Heusler phases M Ru 2 Ga and Ru M 2 Ga (M = first-row transition metal) that are not obvious candidates. Second, how can the search for materials with desired properties be accelerated? One particular application of strong current interest is thermoelectric materials, which present a particular challenge because their optimum performance depends on achieving a balance of many interrelated physical properties. Making use of a recommendation engine developed by Citrine Informatics, we have identified new candidates for thermoelectric materials, including previously unknown compounds (e.g., TiRu 2 Ga with Heusler structure; Mn(Ru 0.4 Ge 0.6 ) with CsCl-type structure) and previously reported compounds but counterintuitive candidates (e.g., Gd 12 Co 5 Bi). An important lesson in these investigations is that the machine-learning models are only as good as the experimental data used to develop them. Thus, experimental work will continue to be necessary to improve the predictions made by machine learning.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.002
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Review · Consensus signal: none
Teacher disagreement score0.002
Threshold uncertainty score0.006

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.002
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.002
Science and technology studies0.0000.001
Scholarly communication0.0010.002
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.194
GPT teacher head0.375
Teacher spread0.181 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designTheoretical or conceptual
Domainnot available
GenreReview

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations148
Published2017
Admission routes2
Has abstractyes

Explore more

Same venueAccounts of Chemical ResearchSame topicMachine Learning in Materials ScienceFrench-language works237,207