rAAV Manufacturing: The Challenges of Soft Sensing during Upstream Processing
Bibliographic record
Abstract
Recombinant adeno-associated virus (rAAV) is the most effective viral vector technology for directly translating the genomic revolution into medicinal therapies. However, the manufacturing of rAAV viral vectors remains challenging in the upstream processing with low rAAV yield in large-scale production and high cost, limiting the generalization of rAAV-based treatments. This situation can be improved by real-time monitoring of critical process parameters (CPP) that affect critical quality attributes (CQA). To achieve this aim, soft sensing combined with predictive modeling is an important strategy that can be used for optimizing the upstream process of rAAV production by monitoring critical process variables in real time. However, the development of soft sensors for rAAV production as a fast and low-cost monitoring approach is not an easy task. This review article describes four challenges and critically discusses the possible solutions that can enable the application of soft sensors for rAAV production monitoring. The challenges from a data scientist's perspective are (i) a predictor variable (soft-sensor inputs) set without AAV viral titer, (ii) multi-step forecasting, (iii) multiple process phases, and (iv) soft-sensor development composed of the mechanistic model.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".