Aligning a Household-Level Service Array Through a Jurisdiction-Wide Child Maltreatment Prevention Effort: Protocol for a Geospatial and Counterfactual Modeling Study
Bibliographic record
Abstract
Background: Child maltreatment is associated with multiple negative outcomes at the individual and societal levels. Children experiencing maltreatment are at greater risk of a host of negative outcomes (eg, psychological disorders, substance use, violent delinquency, suicidality, and adverse educational outcomes). Objective: This study aims to prevent and ameliorate child maltreatment by using a combination of geospatial smoothing via a risk terrain modeling (RTM) framework and counterfactual modeling to identify risky areas and determine the optimal (re)allocation of services to maximally improve maltreatment outcomes. Methods: A 3-stage process is proposed that can iteratively be applied within a collaborating jurisdiction to enable responsive and sustained achievement of identified child welfare outcomes. This process makes use of 2 analytic approaches: geospatial smoothing via an RTM framework and counterfactual modeling. RTM is a spatial analytic approach that uses spatial machine learning methods to estimate the risk of maltreatment based on previous cases of maltreatment and risk factors of the built environment provided by the participating jurisdiction. Using previously validated cases of maltreatment as our target variable (eg, substantiated claims of abuse and neglect) and violent crime data and built environment data as our primary predictor variables, we estimate a series of machine learning models to geospatially smooth the historically identified places at increased risk of child maltreatment. Areas identified as higher risk receive extensive services associated with preventing or limiting child maltreatment, such as prenatal or postnatal care, subsidized daycare, and parental counseling. We make use of counterfactual explanation modeling to optimally align service allocation to maximally improve maltreatment outcomes for future service allocations within a collaborating jurisdiction. The technique leverages a statistical model associating household-level information with maltreatment outcomes to explore combinations of services that would be predicted to achieve optimal and practical recommendations for future service allocation efforts. Constraints can be introduced to this logic, such as service availability and cost. Algorithmic fairness is also a potential consideration during aggregation, with possibilities for both measuring and balancing metrics such as "recourse fairness." Results: As of September 2025, a participating jurisdiction is being recruited. Conclusions: This protocol sets forth a novel approach for the allocation of supportive services for families at risk of child maltreatment through geospatial smoothing via an RTM framework and the maximization of service impact through a counterfactual explanation model. Child maltreatment is an unfortunate and ubiquitous issue in the United States. This proposal builds on jurisdiction-wide public health strategies to allocate services in a data-informed fashion and further align future iterations of the allocation strategy using outcomes-based counterfactual modeling at the household level. The flexibility of the proposed methodology enables its application regardless of the collaborating jurisdiction's preferences and constraints.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.037 | 0.067 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.004 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.003 | 0.005 |
| Insufficient payload (model declined to judge) | 0.055 | 0.008 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".