The clinical performance and cost-effectiveness of two psychosocial assessment models in maternity care: The Perinatal Integrated Psychosocial Assessment study
Bibliographic record
Abstract
PROBLEM: Although perinatal universal depression and psychosocial assessment is recommended in Australia, its clinical performance and cost-effectiveness remain uncertain. AIM: To compare the performance and cost-effectiveness of two models of psychosocial assessment: Usual-Care and Perinatal Integrated Psychosocial Assessment (PIPA). METHODS: Women attending their first antenatal visit were prospectively recruited to this cohort study. Endorsement of significant depressive symptoms or psychosocial risk generated an 'at-risk' flag identifying those needing referral to the Triage Committee. Based on its detailed algorithm, a higher threshold of risk was required to trigger the 'at-risk' flag for PIPA than for Usual-Care. Each model's performance was evaluated using the midwife's agreement with the 'at-risk' flag as the reference standard. Cost-effectiveness was limited to the identification of True Positive and False Positive cases. Staffing costs associated with administering each screening model were quantified using a bottom-up time-in-motion approach. FINDINGS: Both models performed well at identifying 'at-risk' women (sensitivity: Usual-Care 0.82 versus PIPA 0.78). However, the PIPA model was more effective at eliminating False Positives and correctly identifying 'at-risk' women (Positive Predictive Value: PIPA 0.69 versus Usual Care 0.41). PIPA was associated with small incremental savings for both True Positives detected and False Positives averted. DISCUSSION: Overall PIPA performed better than Usual-Care as a psychosocial screening model and was a cost-saving and relatively effective approach for detecting True Positives and averting False Positives. These initial findings warrant evaluation of longer-term costs and outcomes of women identified by the models as 'at-risk' and 'not at-risk' of perinatal psychosocial morbidity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".