Metal-Organic Framework CH4@65bar Adsorbate Probability Distributions
Bibliographic record
Abstract
Metal-organic framework adsorbate probability distributions (APDs) of CH4 at conditions of 65 bar and 298 K. These APDs were used to train ML models in the paper . The training set and development set come from the ARC-MOF database (primarily hypothetical MOFs), while the test set comes from the MOSAEC-DB (exclusively experimentally characterized MOFs from the CSD). All APDs were obtained using GCMC, with framework parameters taken from UFF. The tarballs in this repository contain the following information: CH4_65bar_train.tar.gz -- APDs used in the training set in VASP CHGCAR format. CH4_65bar_dev.tar.gz -- APDs used in the development set in VASP CHGCAR format. CH4_65bar_test_MOSAECDB -- APDs of the MOSAEC-DB test set. CH4_65bar_dev_binding_sites.tar.gz -- Binding sites extracted from the APDs in the CH4_65bar_dev.tar.gz set using our binding site extraction code (GALA). CH4_65bar_test_binding_sites.tar.gz -- Binding sites extracted from the APDs in the CH4_65bar_test.tar.gz set using our binding site extraction code (GALA). The binding_sites.tar.gz files contain one directory per MOF, with the following files/subdirectories in each directory: (Note that "O" represents the adsorbate "CH4", since it is modeled as a single-site, and is following a temporary naming convention in GALA. It does not correspond to an oxygen atom, as the name would imply. "CM" is the centre-of-mass pseudoatom.) Prob_Guest_O_Site_CM_folded.cube -- the APD of the MOF in .cube format. FIELD -- a file in DLPOLY FIELD format defining interaction parameters for guest-host binding energy calculations. GALA.inp -- file containing GALA binding site extraction parameters. gala.log -- GALA log file. gala.out -- GALA output file. slurm-* -- standard output. DL_poly_BS -- a directory generated by GALA to compute the guest-host interaction energy of each binding site. Each subdirectory within DL_poly_BS contains DLPOLY input/output files for each configuration (i.e., the framework, and each binding site configuration). The _ subdirectory contains the framework with a randomly placed guest configuration. This configuration is important to get the electrostatic energy of the framework, which is subtracted from the binding site configuration energies in GALA (the charges of the guest, if any, are zeroed out in this configuration). See the GALA documentation for more details on binding energy calculations. GALA_Output -- a directory generated by GALA containing the following files: O_binding_sites.cif -- all binding sites in .cif format, with the binding sites listed in the .cif file in order of decreasing occupancy. O_binding_sites_fractional.xyz -- fractional coordinates of all binding sites in the MOF O_binding_sites_optimized.cif -- all binding sites optimized using MD as implemented in DLPOLY. In this work, a time-step of 0 ps was used, so this is the same as O_binding_sites.cif. O_gala_binding_sites.xyz -- the Cartesian xyz coordinates of each binding site, ordered in order of decreasing occupancy. This file also contains the guest-host interaction energy of each site (Ebind), the % of Ebind arising from electrostatics (esp%), the van der Waals interaction energy (Evdw), the electrostatic interaction energy (Eesp), and the relative occupancy in % (occ). O_guest_information.xyz -- a file containing adsorbate information, and binding sites listed in order of increasing occupancy, with the absolute occupancy values, fractional, and Cartesian coordinates of each binding site.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.021 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".