Digital Descriptors in Predicting Catalysis Reaction Efficiency and Selectivity
Bibliographic record
Abstract
Accurately controlling the interactions and dynamic changes between multiple active sites (e.g., metals, vacancies, and lone pairs of heteroatoms) to achieve efficient catalytic performance is a key issue and challenge in the design of complex catalytic reactions involving 2D metal-supported catalysts, metal-zeolites, metal-organic catalysts, and metalloenzymes. With the aid of machine learning (ML), descriptors play a central role in optimizing the electrochemical performance of catalysts, elucidating the essence of catalytic activity, and predicting more efficient catalysts, thereby avoiding time-consuming trial-and-error processes. Three kinds of descriptors─active center descriptors, interfacial descriptors, and reaction pathway descriptors─are crucial for understanding and designing metal-supported catalysts. Specifically, vacancies, as active sites, synergize with metals to significantly promote the reduction reactions of energy-relevant small molecules. By combining some physical descriptors, interpretable descriptors can be constructed to evaluate catalytic performance. Future development of descriptors and ML models faces the challenge of constructing descriptors for vacancies in multicatalysis systems to rationally design the activity, selectivity, and stability of catalysts. Utilization of generative artificial intelligence and multimodal ML to automatically extract descriptors would accelerate the exploration of dynamic reaction mechanisms. The transferable descriptors from metal-supported catalysts to artificial metalloenzymes provide innovative solutions for energy conversion and environmental protection.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".