NMR Crystallography Structure Determinations with <sup>1</sup>H Chemical Shifts. GIPAW DFT Calculation Quality Can Be Substantially Degraded, but Nearly Identical Outputs Relative to Benchmark Computations Are Obtained: Why and So What?
Bibliographic record
Abstract
Nuclear magnetic resonance (NMR) crystallography may be used in various solid-state structural characterization tasks. For organic compounds in this context, proton isotropic chemical shifts [δ iso ( 1 H)] are routinely used. It is typical to pair experimentally measured proton δ iso values with δ iso values that were computationally generated from crystal structure models. This can yield a δ iso ( 1 H) root-mean-squared deviation (RMSD) value for each crystal structure model. In this study, we monitor the way in which gauge including projector augmented wave (GIPAW) density functional theory (DFT) computations of 1 H δ iso values can be influenced by the quality of the computational input parameters. We consider 126 computationally generated (using crystal structure prediction, CSP) crystal structures for three molecules: cocaine (30 structures), flutamide (21 structures), and ampicillin (75 structures). The quality parameters selected are the plane wave energy cutoff ( E cut ), and the k -point grid used to sample reciprocal (i.e., momentum) space. We also probe the utility of performing one-parameter and two-parameter linear mappings for transforming computed hydrogen isotropic magnetic shielding values (σ iso ) into computed δ iso ( 1 H) values. We find that both E cut and the k -point grid can be degraded substantially (e.g., E cut ∼ 25 Ry) and yet still produce very similar computed δ iso ( 1 H) values. We consider the mechanisms under GIPAW DFT that contribute to computed hydrogen σ iso values to help understand this robustness: many contributions are zero or cancel out when converting σ iso values to δ iso ( 1 H) values via the linear mapping. The robust nature of computed δ iso ( 1 H) values leads to consistent estimates of δ iso ( 1 H) RMSD values. It is then demonstrated using cocaine and flutamide that when δ iso ( 1 H) RMSD values are used in NMR crystallography tasks such as structure selection/determination, the quality of the GIPAW DFT computation can be severely degraded and still produce identical outcomes to those that used a more computationally intensive protocol. Ampicillin is selected as a practical example to probe how our findings might reasonably be applied in the structure determination of a complex organic molecule. We propose that relatively modest quality GIPAW DFT computations (i.e., E cut = 35 Ry and a 1 × 1 × 1 k -point grid) may be used to first filter out obviously poor structure candidates. Subsequently, slightly higher quality GIPAW DFT computations can be used for structure selection/determination. Our findings indicate that it should be possible to, on average, reduce the computational resources required in such NMR crystallography tasks by approximately a factor of 3–4 in terms of CPU time.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".