MétaCan
Menu
Back to cohort
Record W7115819204

Rapid Model-Based Recipe Design from Limited-Sample Datasets: Experimental Validation in Nanoparticle and Microgel Systems

2025· dissertation· en· W7115819204 on OpenAlexfundno aff

Bibliographic record

VenueMacSphere (McMaster University) · 2025
Typedissertation
Languageen
FieldComputer Science
TopicComputational Drug Discovery Methods
Canadian institutionsnot available
FundersNatural Sciences and Engineering Research Council of CanadaCanada Research ChairsMcMaster University
KeywordsReliability (semiconductor)MulticollinearityMetric (unit)Partial least squares regressionProjection (relational algebra)Latent variableInverseCluster analysisNonlinear system
DOInot available

Abstract

fetched live from OpenAlex

In data-constrained experimental domains such as nanoparticle engineering, microgel synthesis, and pharmaceutical formulation, researchers frequently face the challenge of modeling systems governed by complex and highly nonlinear relationships among variables. These applications often involve limited datasets due to the cost, time, and resource demands of generating new samples, making conventional trial-and-error approaches inefficient. As a result, there is a growing need for data-driven methodologies that can reliably predict product behavior and guide recipe design using minimal experimental input. This thesis presents a series of strategies to enhance prediction accuracy and reliability in small datasets through localized modeling, quantitative model reliability assessment, and guided expansion of the available dataset. The first contribution involves coupling Latent Variable Modeling (LVM) with clustering to create local Partial Least Squares (PLS) models tailored to subsets of similar samples. This combination simplifies the underlying data structure by reducing multicollinearity via projection into latent space and grouping structurally similar data, thereby improving prediction fidelity. The framework was validated using the prediction of the Volume Phase Transition Temperature (VPTT) of dual-responsive microgels—a property influenced by several formulation variables—with results that showed significantly improved prediction accuracy. Building on these advances, the second contribution focuses on the inverse problem of design space identification—determining input configurations that are most likely to yield desired output properties. To do this robustly, the Prediction Reliability Enhancing Parameter (PREP) is introduced, a novel metric that unifies multiple LVM alignment diagnostics including Hotelling T², Squared Prediction Error (SPE), and score alignment factors into a single predictive reliability score. PREP is calibrated in a data-driven, case-specific manner and facilitates the ranking of candidate formulations by their expected predictive reliability. Extensive validation on simulated datasets has demonstrated that PREP significantly accelerates the identification of optimal solutions, particularly under highly nonlinear conditions and limited data regimes. PREP was subsequently deployed across real experimental case studies involving the formulation of nanoparticles and microgels. In one study, a microgel with a tightly constrained particle size of ~100 nm was successfully designed from an initial dataset spanning sizes of 170–900 nm, with PREP delivering a near-target solution in minimal iterations while competing design approaches failed under the applied constraints. In another case, PREP enabled the identification of polyelectrolyte complexes with particle sizes below 200 nm and polydispersity indices under 0.2, again demonstrating superior efficiency and accuracy relative to conventional approaches. Overall, this work offers a practical and scalable pathway for predictive modeling and recipe design in settings constrained by data scarcity and high experimental costs. The methodologies developed—particularly the integration of local LVM models and the PREP-based design space identification—can be broadly applied to other high-value domains requiring precision formulation and optimization such as drug delivery, nanomedicine, and advanced materials development.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.009
metaresearch head score (Gemma)0.015
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.009
Threshold uncertainty score0.048

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0090.015
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0010.001
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.039
GPT teacher head0.260
Teacher spread0.222 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueMacSphere (McMaster University)Same topicComputational Drug Discovery MethodsFrench-language works237,207