Scalable methods for modelling complex biochemical networks
Bibliographic record
Abstract
In cells, complex networks of interacting biomolecules process both environmental and endogenous signals to control gene expression and other cellular processes. This poses a challenge to researchers who attempt to develop mathematical and computational models of biochemical networks that reflect this complexity. In this thesis, I propose methods that help manage complexity by exploiting the finding that, as for other biological systems, cellular networks are characterized by a modularity that appears at all levels of organization.The first part of this work focuses on the modular properties of proteins and how their function can be characterized through their structure and allosteric properties. I develop a modular rule-based framework and formal modelling language that describes the computations performed by allosteric proteins and that is rooted in biophysical principles. Rule-based modelling conventionally addresses the problem of combinatorial complexity, whereby protein interactions can generate a combinatorial explosion of protein complex states. However, I explore how these same interactions can potentially require a combinatorial number of parameters to describe them. I demonstrate that my rule-based framework effectively addresses this problem of regulatory complexity, and describes allosteric proteins and networks in a unified, consistent, and modular fashion. I use the framework in three applications. First, I show that allostery can make macromolecular assembly more efficacious when a protein that joins two separable parts of a complex is present in excessively high concentrations. Second, I demonstrate that I can straightforwardly analyze the complex cooperative interactions that arise when competitive ligands bind to a multimeric protein. Third, I analyze a new model of G protein-coupled receptor signalling and demonstrate that it explains the functional selectivity of these receptors while being parsimonious in the number of parameters used. Overall, I find that my rule-based modelling framework, implemented as the Allosteric Network Compiler software tool, can ease of modelling and analysis of complex allosteric interactions.If cellular networks are modular, this implies that small sub-systems can be studied in isolation, provided that external inputs and perturbations to the system can be modelled appropriately. However, cellular networks are subject to both intrinsic noise, which is endogenous to the system, but also extrinsic noise, arising from noisy inputs. Furthermore, many inputs may be dynamic, whether due to experimental protocols or perhaps reflecting the cyclic process of cell division. This motivates my development, in the second part of this work, of efficient stochastic simulation algorithms for biochemical networks that can accommodate time-varying biochemical parameters. Starting from Gillespie's well-known First Reaction Method and Gibson and Bruck's Next Reaction Method, I develop two new algorithms that allow time-varying inputs of arbitrary functional form while scaling well to systems comprising many biochemical reactions. I analyze their scaling properties and find that a modified First Reaction Method may scale better than a modified Next Reaction Method in some applications.The third and last part of this thesis introduces a new software tool, Facile, that eases the creation, update and simulation of biochemical network models. Models created through a simple and intuitive textual language are automatically converted into a form usable by downstream tools, for example ordinary differential equations for simulation by Matlab. Also, Facile conveniently accommodates mathematical and time-varying expressions in rate laws.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".