Breaking barriers in surface tension prediction of aqueous organics: A Hansen parameter-based machine learning model optimized with a genetic algorithm
Bibliographic record
Abstract
A practical and generalizable machine learning model was developed to predict the surface tension of aqueous organic solutions, relevant to a wide range of applications, including CO 2 capture, pharmaceuticals, membranes, and energy. Aqueous solutions of various organic compounds—such as alcohols, acids, amines, surfactants, and sophorolipids—were used to train a multilayer perceptron artificial neural network (MLPANN), a powerful machine learning tool. To balance simplicity and accuracy, the network’s architecture was optimized using a genetic algorithm, rather than relying on traditional or non-traditional calculation methods. Furthermore, to enhance practicality, in addition to essential variables such as temperature, composition, and molecular weight, only the Hansen solubility parameters (HSPs) of water and the organic component were used as inputs. The proposed approach successfully correlated 319 training data points, yielding an average absolute relative deviation (AARD) of 2.61 %. In prediction mode, it achieved an AARD of 3.53 % for 66 testing data points, demonstrating robust predictive accuracy. Its extrapolation capability was further validated on the unseen monoethanolamine (MEA) + water system, where it achieved an AARD of 2.41 % across a broad range of temperatures and compositions. This approach presents a practical and novel solution for predicting the surface tension of any aqueous organic solution, striking an optimal balance between simplicity, accuracy, and generality. Its capacity to handle both pure and binary mixtures with minimal input makes it easily extendable to ternary and multicomponent systems. Furthermore, it offers valuable insights for the modeling of other thermophysical properties.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".