Harnessing the power of ab initio calculations, distributed computing and machine learning to efficiently locate extreme molecules for use in carbon-based solar cells (Final Technical Report)
Bibliographic record
Abstract
The use of high-throughput virtual screening (HTVS) tools is a powerful tool to expedite the materials discovery of commercially relevant materials. In previous years, our group has developed a molecular discovery platform to generate libraries in order to obtain suitable candidates for different applications, starting from the Harvard Clean Energy Project [1,2]. This platform is suitable to test in-silico on traditional supercomputing clusters and shared resources, for example, in the IBM World Community Grid. In this project, we used the molecular discovery platform to create and screen a library of candidates of organic photovoltaic (OPVs) molecules. Based on a set of candidates created with combinations of molecular moieties, we were able to filter, by conformation stability, the energy of electronic orbitals and approximated power conversion efficiencies (PCE). To improve the predictions of orbital energies calculated and the PCEs, we used Gaussian Process regression and two sets of molecules. These sets correspond to electronic structure calculations of a higher level of theory and experimental PCE values, respectively. Finally, we selected a subset of the best candidates (molecules with a PCE higher than 10%) to understand its absorbance properties with TD-DFT. This project has demonstrated the capabilities of our molecular discovery platform for HTVS. Finally, machine learning can help us to introduce more complex effects included in bulk conditions and computational intensive calculations on models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".