Rapid Prediction of Solvation Free Energy. 1. An Extensive Test of Linear Interaction Energy (LIE)
Bibliographic record
Abstract
The present study provides a comprehensive systematic analysis on the applicability of the linear interaction energy (LIE) approximation to the prediction of gas-to-water transfer (hydration) free energy. The study is based on molecular dynamics simulations in explicit solvent for an extensive and diverse hydration data set comprising 564 neutral compounds with measured hydration free energies, including a "traditional" data set and the more challenging drug-like SAMPL1 data set. A highly correlative LIE model was achieved without empirical scaling of the solute-solvent interaction energy terms along with a cavity term calibrated to the experiment. This model was particularly accurate for the "traditional" data set and of acceptable accuracy for the SAMPL1 data set, with mean-unsigned-errors below 1 kcal/mol and slightly above 2 kcal/mol, respectively. We have analyzed the sensitivity of the LIE model to several parameters such as continuum correction terms applied outside the explicit water shell, the impact of various charging methods, the applicability of single-conformer representation of the solute, and the inclusion of internal energy terms. The parameters with the greatest sensitivity are the charging methods used, with AM1BCC-SP (without AM1 geometry optimization) charges favored over AM1BCC-OPT and RESP charges. The inclusion of the change in intramolecular van der Waals and electrostatic energies between the solution and gas phases can also lead to improved prediction accuracies. Functional group based error analysis identified several chemical classes as minor outliers with systematic errors. A direct comparison of the LIE and free energy perturbation (FEP) approaches using the same force field and charging method shows that the LIE approximation is at least as accurate as the FEP approach with a reduction of computing time by at least 1 order of magnitude.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".