Development and Refinement of a Chatbot for Birthing Individuals and Newborn Caregivers: Mixed Methods Study
Bibliographic record
Abstract
BACKGROUND: The 42 days after delivery ("fourth trimester") are a high-risk period for birthing individuals and newborns, especially those who are racially and ethnically marginalized due to structural racism. OBJECTIVE: To fill a gap in the critical "fourth trimester," we developed 2 ruled-based chatbots-one for birthing individuals and one for newborn caregivers-that provided trusted information about postbirth warning signs and newborn care and connected patients with health care providers. METHODS: A total of 4370 individuals received the newborn chatbot outreach between September 1, 2022, and December 31, 2023, and 3497 individuals received the postpartum chatbot outreach between November 16, 2022, and December 31, 2023. We conducted surveys and interviews in English and Spanish to understand the acceptability and usability of the chatbot and identify areas for improvement. We sampled from hospital discharge lists that distributed the chatbot, stratified by prenatal care location, age, type of insurance, and racial and ethnic group. We analyzed quantitative results using descriptive analyses in SPSS (IBM Corp) and qualitative results using deductive coding in Dedoose (SocioCultural Research Consultants). RESULTS: Overall, 2748 (63%) individuals opened the newborn chatbot messaging, and 2244 (64%) individuals opened the postpartum chatbot messaging. A total of 100 patients engaged with the chatbot and provided survey feedback; of those, 40% (n=40) identified as Black, 27% (n=27) identified as Hispanic/Latina, and 18% (n=18) completed the survey in Spanish. Payer distribution was 55% (n=55) for individuals with public insurance, 39% (n=39) for those with commercial insurance, and 2% (n=2) for uninsured individuals. The majority of surveyed participants indicated that chatbot messaging was timely and easy to use (n=80, 80%) and found the reminders to schedule the newborn visit (n=59, 59%) and postpartum visit (n=66, 66%) useful. Across 23 interviews (n=14, 61% Black; n=4, 17% Hispanic/Latina; n=2, 9% in Spanish; n=11, 48% public insurance), 78% (n=18) of interviewees engaged with the chatbot. Interviewees provided positive feedback on usability and content and recommendations for improving the outreach messages. CONCLUSIONS: Chatbots are a promising strategy to reach birthing individuals and newborn caregivers with information about postpartum recovery and newborn care, but intentional outreach and engagement strategies are needed to optimize interaction. Future work should measure the chatbot's impact on health outcomes and reduce disparities.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.046 | 0.035 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.004 | 0.002 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".