Abstract PR-04: A practical framework for operationalizing responsible and equitable AI in healthcare: Tackling bias, inequity, and implementation challenges
Bibliographic record
Abstract
Abstract Background: Artificial intelligence (AI) is promising to rapidly transform healthcare by enhancing clinical workflows and improving patient outcomes. However, the integration of AI solutions also carries significant risk of harm due to discriminatory performance and inequitable outcomes across diverse patient populations. Existing frameworks aimed at promoting responsible AI development, such as SPIRIT-AI, CONSORT-AI, and TRIPOD+AI, provide guidelines for clinical trial design but lack concrete recommendations to identify and mitigate bias during clinical integration. Frameworks emphasizing ethical principles like equity, transparency, and accountability, including HEAAL, JustEFAB, and the Normative Framework, similarly fall short of offering detailed operational guidance for real-world AI deployment. Recognizing these gaps, we developed a novel Framework for Responsible AI Deployment in healthcare settings, incorporating structured, actionable steps to identify, mitigate, and monitor biases throughout the AI lifecycle. Methodology: Our framework (https://github.com/pmcdi/responsible-ai) was developed through a multidisciplinary collaborative approach whereby stakeholders with expertise in biostatistics, machine learning, ethics, clinical care, institutional governance, diversity and inclusion, and patient advocates, synthesized insights from existing frameworks and engaged in iterative and structured feedback sessions to ensure practical applicability and robustness. Results: This framework is organized into four distinct stages: (1) Problem Identification and Study Design, emphasizing equity-focused clinical question formulation and ethical compliance; (2) Model Training and Development, addressing biases in retrospective data and ensuring transparent performance evaluations; (3) Silent Deployment and Clinical Evaluation, prospectively validating model fairness and clinical applicability without direct patient impact; and (4) Clinical Deployment and Lifecycle Monitoring, providing continuous oversight of AI systems integrated into clinical workflows, emphasizing patient and clinician education, compliance monitoring, and adaptive maintenance. The framework is accompanied by a supplemental appendix which contextualizes each stage with concrete detail such as recommended methods, pain points to consider, and academic references for exploration. Conclusions: Our framework addresses critical shortcomings in current practices to facilitate ethical and equitable AI deployment in healthcare. We are actively working with researchers at the Princess Margaret Cancer Centre to evaluate its utility across a breadth of clinical AI solutions at all stages of development. This framework can help institutions meet their ethical obligations; ensure AI-driven innovations align with foundational healthcare principles of fairness, safety, and quality; safeguard against harm; and ultimately improve trust in AI-enhanced clinical care. Citation Format: Benjamin Grant, Mattea Welch, Christopher Deutschman, Clare McElcheran, Adam Badzynski, Jennifer A.H. Bell, Andrew Hope, Robert C. Grant, Tran Truong, Kelly Lane, Patti Leake, Divya Sharma, Ian Stedman, Mike Lovas, Jeremy Petch, Muammar Kabir, Alejandro Berlin, James A. Anderson, Benjamin Haibe-Kains. A practical framework for operationalizing responsible and equitable AI in healthcare: Tackling bias, inequity, and implementation challenges [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr PR-04.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.234 | 0.224 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.006 | 0.003 |
| Science and technology studies | 0.004 | 0.019 |
| Scholarly communication | 0.014 | 0.013 |
| Open science | 0.009 | 0.016 |
| Research integrity | 0.008 | 0.011 |
| Insufficient payload (model declined to judge) | 0.011 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".