The road to virtual chemistry: computer-aided molecular design skirting the boundary between structure and Ligand-based approaches
Bibliographic record
Abstract
Novel synthetic methodologies that are effective, environmentally friendly and efficient are becoming ever increasingly difficult to design. In fact, the development of such techniques is often labor-intensive, wasteful and costly. Several years ago, instrumental techniques, such as nuclear magnetic resonance, high performance liquid chromatography and mass spectrometry, were integrated into the chemistry toolbox and their maturity significantly accelerated the process of synthetic discovery. Surprisingly and contrastingly, computational advances have not yet been incorporated as far as the imagination can take them. Fifty years after Gordon E. Moore, the co-founder of Intel, first predicted a yearly two-fold expansion of computational power, we are attaining its peak, and yet information technologies are still under-utilized in chemical settings. Until now, computational techniques employed in designing chemical and biochemical synthesis have been merely a tease. Expanding the abilities of computational molecular discovery methods is an attractive solution to exploring a vast amount of unknown synthetic approaches. Furthermore, making these virtual methodologies accessible to the organic and medicinal chemistry communities will allow them to reach their full potential. Currently, only a handful of research groups in Canada truly blend computational and organic chemistry and often in a rationalization capacity rather than as a design strategy. This thesis describes efforts to develop new computational design approaches for small molecules and biological structures and apply them to organo- and biocatalytic research programs. A contemporary perspective on software programs is required for their inclusion in the chemistry toolbox for several reasons. Currently, hundreds, if not thousands, of computational chemistry software packages exist; however, in most cases, it is only the developers that make use of these tools, the ones that simulate chemical phenomena, as opposed to visualization software suites. A significant lack of usability – simplicity in running routine experiments – is likely one of the largest causes for this disappointing reality. Accurate results are also necessary to build trust from the experimental chemistry community. These issues were the focus of this work to demonstrate that the integration of computational tools within organic chemistry is not only plausible, but increasingly valuable when no advanced, expert training is necessary.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".