Enhancing Symbolic Regression and Universal Physics-Informed Neural Networks with Dimensional Analysis
Bibliographic record
Abstract
In engineering and applied mathematics, developing accurate mathematical models to predict and understand real-world phenomena is of utmost importance. Symbolic regression is a useful machine learning-based tool to fit models but it can be computationally expensive. We present a new method for enhancing symbolic regression for differential equations via dimensional analysis, specifically the Buckingham $Π$ theorem and Ipsen's method. Since symbolic regression often suffers from high computational costs and overfitting, nondimensionalizing datasets reduces the number of input variables, simplifies the search space, and ensures that derived equations are physically meaningful. As a first step, we combine dimensional analysis with the PySR symbolic regression algorithm to show that dimensional analysis improves the accuracy of recovering algebraic equations. The results demonstrate that transforming data into a dimensionless form significantly improves the training and test error of the symbolic expressions found. Then, as our main contribution, we perform nondimensionalization guided by Ipsen's method. We then incorporate the nondimensionalized equation into a pipeline combining Universal Physics-Informed Neural Networks and symbolic regression to recover the unknown term when a differential equation is only partially known. We find that symbolic regression is able to better recover the unknown term after nondimensionalizing the data, under both noisy and noiseless conditions. These findings suggest that integrating dimensional analysis with symbolic regression can significantly lower computational costs and increase accuracy, providing a robust framework for automated discovery of governing equations in complex systems when data is limited.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".