Machine Learning–Guided Structure–Activity Discovery of Polymer Configurations in Lipid Nanoparticles for Kiss-and-Run Endosomal Escape
Bibliographic record
Abstract
Abstract Endosomal escape remains a major barrier to effective nucleic acid delivery via lipid nanoparticles (LNPs). Here, we address this challenge by incorporating a pH-sensitive polymer, polyhistidine, into LNPs (pLNPs) to facilitate endosomal escape, with a focus on optimizing the polymer’s molecular weight (MW) and configuration—parameters that remain largely unexplored. Through systematic engineering, we designed linear and branched polyhistidine architectures with varied MWs and configurations. In vivo screening identified an optimized pLNP formulation incorporating a symmetrical bis-lysine histidine dendron with a MW of ∼1800 g/mol, which achieved a 266-fold increase in liver bioluminescence following intravenous delivery of luciferase mRNA compared to standard LNPs at an equivalent RNA dose. Mechanistic studies revealed that polymer configuration within pLNPs is critical for eliciting the proton sponge effect, leading to osmotic swelling and endosomal rupture. This configuration also promoted rapid endosomal membrane destabilization via a kiss-and-run mechanism, enabling efficient cytosolic release. When delivering base editor mRNA and single-guide RNA, the optimized pLNPs achieved 8% gene editing efficiency in the mouse liver at a low dose of 0.1 mg/kg, compared to 1% with standard LNPs. To accelerate discovery and address macromolecular design challenges, we developed a machine learning (ML) framework based on amino acid-level graph neural networks (GNNs). This approach identified branched, dendritic configurations with densely arranged histidine residues on a multivalent core as key determinants of delivery performance. The top ML-predicted candidate, NS535, achieved a 705-fold increase in liver bioluminescence over standard LNPs, validating our data-driven design strategy. Together, these findings establish a closed-loop platform integrating rational design, mechanistic validation, and ML-guided optimization to advance RNA delivery. By elucidating structure-activity relationships for polyhistidine carriers and demonstrating efficient, low-dose genome editing, this work provides a blueprint for next-generation nucleic acid therapeutics.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".