MétaCan
Menu
Back to cohort
Record W4412626405 · doi:10.2196/80416

A Multiassessment and Multiprofessional Agents Approach for Medical Chatbot Risk Estimation: Development and Evaluation Study

2025· preprint· en· W4412626405 on OpenAlexvenueno aff
Lenard Paulo V. Tamayo, Tomohiro Nishiyama, Shaowen Peng, Shoko Wakamiya, Eiji Aramaki

Bibliographic record

VenueJMIR Medical Informatics · 2025
Typepreprint
Languageen
FieldMedicine
TopicArtificial Intelligence in Healthcare and Education
Canadian institutionsnot available
Fundersnot available
KeywordsPreprintChatbotEstimationComputer scienceData scienceArtificial intelligenceEngineeringWorld Wide WebSystems engineering

Abstract

fetched live from OpenAlex

BACKGROUND Assessing chatbot responses across 3 domains—medical, ethical, and legal—is essential to ensuring the safe use of artificial intelligence in health care. Although advancements in the use of large language models (LLMs) show significant improvements in evaluating question-answer datasets, such as multiple-choice medical exams, existing systems use general LLMs without incorporating specialized domain knowledge. They rely on standardized instructions without integrating real-world information, and ensemble methods such as majority voting fail to resolve disagreements among agents, resulting in misclassification and challenges in risk assessment. OBJECTIVE This study aims to design, develop, and evaluate a synergistic approach for assessing risks associated with chatbot responses using multiassessment (MA) and multiprofessional agents (MPAs). METHODS We designed and developed an approach consisting of MA and MPA, specifically initial assessment (MA1), which internalizes 3 roles and provides an initial risk estimation, and final assessment (MA3), which aims to reach a final consensus based on the previous assessments (MA1 and MA2), with each using 1 LLM. The verification assessment (MA2) incorporates an MPA or role-based LLM specialized agents for each risk domain (medical, ethical, and legal). We evaluated the proposed approach using the MedNLP-CHAT (Medical Natural Language Processing for AI Chat) corpus (N=226; 100 train, 126 test), covering baseline, enhanced prompt, embedding-based search, and retrieval-augmented generation (RAG). Primary metrics included macro F1-score and joint accuracy to evaluate system performance, along with CI and paired macro F1-score difference (Δ) as supporting metrics to assess the approach’s effectiveness. RESULTS The MA-MPA framework integrated with RAG achieved the highest average macro F1-score of 0.800 across risk domains and a joint accuracy of 76 (60.3%) correct predictions across all risk domains out of 126 question-answer pairs, with notable improvements over the best reported Eighteenth NII Testbeds and Community for Information Access Research Project (NTCIR-18) MedNLP-CHAT systems in the ethical (+0.252) and legal (+0.096) risk domains, while the medical domain showed a modest increase of +0.070. The MA approach contributed the largest gains, particularly from MA1 to MA2, with paired macro F1-score gains ranging from +0.176 to +0.214 across systems. The MPA approach performed better when integrated with MA and external knowledge, with paired bootstrap estimates showing a gain of +0.037 (95% CI 0.003-0.074) over baseline; however, joint accuracy gains were not evident (95% CI –2.9% to 7.7%), and gains relative to the enhanced prompt were small. Notably, MA alone achieved higher joint accuracy than RAG (62.7% vs 60.3%), indicating a metric-specific trade-off rather than consistent superiority across all metrics. CONCLUSIONS The MA-MPA approach shows potential for improving risk estimation in chatbot responses. The results suggest that the framework is particularly useful for enhancing balanced overall performance, especially when combined with external knowledge, although the medical risk domain remains challenging. Furthermore, more specialized LLMs may further improve contextually grounded risk estimation.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.015
metaresearch head score (Gemma)0.030
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.015
Threshold uncertainty score0.080

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0150.030
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.001
Science and technology studies0.0010.001
Scholarly communication0.0020.003
Open science0.0020.003
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0040.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.237
GPT teacher head0.538
Teacher spread0.301 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Medical InformaticsSame topicArtificial Intelligence in Healthcare and EducationFrench-language works237,207