Governing AI in Mental Health: 50-State Legislative Review
Bibliographic record
Abstract
Background: Mental health-related artificial intelligence (MH-AI) systems are proliferating across consumer and clinical contexts, outpacing regulatory frameworks and raising urgent questions about safety, accountability, and clinical integration. Reports of adverse events, including instances of self-harm and harmful clinical advice, highlight the risks of deploying such tools without clear standards and oversight. Federal authority over MH-AI is fragmented, leaving state legislatures to serve as de facto laboratories for MH-AI policy. Some states have been highly active in this area during recent legislative sessions. Yet, clinicians and professional organizations have mainly remained absent or sidelined from public commentary and policymaking bodies, raising concerns that new laws may diverge from the realities of mental health care. Objective: To systematically analyze recent state-level legislation relevant to MH-AI, categorize bills by relevance to mental health, identify major regulatory themes and gaps, and evaluate implications for clinicians and patients. Methods: We conducted a systematic analysis of bills introduced in all 50 US states between January 1, 2022, and May 19, 2025, using standardized searches on the legislative research website (LegiScan). Bills were screened and categorized using a custom 4-tier taxonomy based on their applicability to MH-AI. Bills passing threshold review were coded by topic using a 25-tag system developed through iterative consensus. Legally trained reviewers adjudicated final classifications to ensure consistency and rigor. Results: Among 793 state bills reviewed, 143 were identified as potentially impactful to MH-AI: 28 explicitly referenced mental health uses, while 115 had substantial or indirect implications. Of these 143 bills, 20 were enacted across 11 states. Legislative efforts varied widely, but 4 thematic domains consistently emerged: (1) professional oversight, including deployer liability and licensure obligations; (2) harm prevention, encompassing safety protocols, malpractice exposure, and risk stratification frameworks; (3) patient autonomy, particularly in areas of disclosure, consent, and transparency; and (4) data governance, with notable gaps in privacy protections for sensitive mental health data. Conclusions: State legislatures are rapidly shaping the regulatory landscape for MH-AI, but most laws treat mental health as incidental to broader artificial intelligence or health care regulation. Explicit mental health provisions remain rare, and clinician and patient perspectives are seldom incorporated into policymaking. The result is a fragmented and uneven environment that risks leaving patients unprotected and clinicians overburdened. Mental health professionals must proactively engage with legislators, professional organizations, and patient advocates to ensure that emerging frameworks address oversight, harm, autonomy, and privacy in ways that are clinically realistic, ethically sound, and supportive of flexible-but responsible-innovation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.099 | 0.253 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.004 |
| Bibliometrics | 0.034 | 0.032 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.006 | 0.007 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".