Assessment of Density Functional Theory Methods for the Structural Prediction of Transition and Post-Transition Metal–Nucleic Acid Complexes
Bibliographic record
Abstract
Understanding the structure of metal–nucleic acid systems is important for many applications such as the design of new pharmaceuticals, metal detection platforms, and nanomaterials. Herein, we explore the ability of 20 density functional theory (DFT) functionals to reproduce the crystal structure geometry of transition and post-transition metal–nucleic acid complexes identified in the Protein Data Bank and Cambridge Structural Database. The environmental extremes of the gas phase and implicit water were considered, and analysis focused on the global and inner coordination geometry, including the coordination distances. Although gas-phase calculations were unable to describe the structure of 12 out of the 53 complexes in our test set regardless of the DFT functional considered, accounting for the broader environment through implicit solvation or constraining the model truncation points to crystallographic coordinates generally afforded agreement with the experimental structure, suggesting that functional performance for these systems is likely due to the models rather than the methods. For the remaining 41 complexes, our results show that the reliability of functionals depends on the metal identity, with the magnitude of error varying across the periodic table. Furthermore, minimal changes in the geometries of these metal–nucleic acid complexes occur upon use of the Stuttgart–Dresden effective core potential and/or inclusion of an implicit water environment. The overall top three performing functionals are ωB97X-V, ωB97X-D3(BJ), and MN15, which reliably describe the structure of a broad range of metal–nucleic acid systems. Other suitable functionals include MN15-L, which is a cheaper alternative to MN15, and PBEh-3c, which is commonly used in QM/MM calculations of biomolecules. In fact, these five methods were the only functionals tested to reproduce the coordination sphere of Cu 2+ -containing complexes. For metal–nucleic acid systems that do not contain Cu 2+, ωB97X and ωB97X-D are also suitable choices. These top-performing methods can be utilized in future investigations of diverse metal–nucleic acid complexes of relevance to biology and material science.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".