A test of various computational solvation models on a set of “difficult” organic compounds
Bibliographic record
Abstract
Various dielectric continuum models in Gaussian 03, based on the SCRF approach, PCM, CPCM, DPCM, IEFPCM, IPCM, and SCIPCM, have been tested on a set of 54 highly polar, generally polyfunctional compounds for which experimental solvation energies are available. These compounds span a range of 13 kcal/mol in ΔG t . The root-mean-square (RMS) errors for the full set of compounds range from 2.48 for DPCM to 1.77 for IPCM. For each method, classes of compounds which were not handled well could be identified. If these classes of compounds were omitted, the performance improved, and ranged from 1.58 (PCM, 39 compounds) to 1.02 (IPCM, 42 compounds). Models in the PCM family (PCM, CPCM, DPCM, and IEFPCM) with the recommended UAHF or UAKS sets of radii rely on a highly parameterized definition of the solvent cavity. Where this parameterization was inadequate, the calculated solvation energies were less reliable. This has been demonstrated by devising a new parameterization for PCM and halogen compounds, which markedly improves performance for polyhalogen compounds. The effective radius for the portion of the cavity centered on a halogen atom was assumed to be linear in the electron-withdrawing or -donating properties of the rest of the molecule as measured by Hammett σ (for halogens on aromatic rings) or Taft σ* (for halogens on aliphatic carbons). This new parameterization for PCM was tested on a set of 45 aliphatic and 22 aromatic polyhalogen compounds and shown to do well. IPCM, which was already the best of the methods in Gaussian, can be considerably improved by a parameterization to allow for cavitation, dispersion, and hydrogen bonding. A large set of compounds was used for the parameterization to have multiple examples for each parameter and as far as possible to have molecules with multiple instances of each structural feature. In the end, 15 parameters were found to be defined by the data for 241 compounds. With this parameter set, the RMS error for the set used for fitting was 0.81 kcal/mol, and the RMS error for the original set of 54 compounds was 0.85. With this new parameterization, IPCM is clearly the best of the methods available in Gaussian 03.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".