MétaCan
Menu
Retour à la cohorte
Enregistrement W4402582511 · doi:10.1097/j.pain.0000000000003389

AI and the ethics of techno-solutionism in pain management

2024· article· en· W4402582511 sur OpenAlexafffund
Daniel Z. Buchman

Notice bibliographique

RevuePain · 2024
Typearticle
Langueen
DomaineSocial Sciences
ThématiqueDiversity and Career in Medicine
Établissements canadiensUniversity of TorontoCentre for Addiction and Mental HealthUniversity Health Network
Organismes subventionnairesInstitute of Neurosciences, Mental Health and AddictionCanadian Institutes of Health Research
Mots-clésPain managementMedicinePsychologyBusinessAnesthesia

Résumé

récupéré en direct d'OpenAlex

Young et al29 examined 2 large language models' (LLMs) pain management recommendations for the 4 most common reasons for pain-related emergency department visits. Their goal was to understand if LLMs recommendations for opioid treatment vary based on patient race/ethnicity and sex, and whether LLMs eliminate or worsen biases. Using 40 real patient case descriptions, and omitting patients' sex, race, ethnicity, and current medication use from the model, the authors instructed LLMs (GPT-4 and Gemini) to provide subjective pain ratings (severity and rating from 1 to 10) and pharmacologic treatment recommendations. The authors found that although there were discrepancies between LLMs, they did not show preferential opioid treatment for one group over another based on combinations of race/ethnicity or sex. Although Young et al.29 note that there is evidence that artificial intelligence (AI) systems like LLMs can exacerbate race-based inequities, they suggest that LLMs may help mitigate clinician bias and support equitable pain management. I commend the authors for tackling this critical issue. However, I argue that the appeal of using objective AI-based solutions like LLMs to overcome complex socio-structural injustices in pain management reflects a techno-solutionism that will likely not remedy the problems it seeks to solve and could have unintended ethical consequences. There is widespread enthusiasm that AI systems and precision medicine generally will improve healthcare decision making. Although research on state-of-the-art LLMs demonstrate that they can reproduce several tasks, the evidence for clinical impact remains limited.12 Part of the motivation to use AI systems in pain management is to overcome the perceived limitations associated with uncertainty, subjectivity, and invisibility of pain in a medical culture that prioritizes certainty, objectivity, and visibility.4,5,9 For example, researchers are combining neuroimaging with forms of AI such as machine learning to identify a brain-based biomarker of chronic pain.7,22,27 These research programs are based on assumptions of mechanical objectivity, where objectivity is achieved with standardized methods, mechanical or automated processes, and instruments that attempt to minimize human bias and subjectivity. The outcomes are therefore considered more reliable and unbiased.6,9,25 The idea that LLMs could potentially fix clinician biases and inequities in pain management related to sexism, racism, and other systems of oppression is understandably appealing. Despite bias being a massive problem in pain management, LLMs potential use in this context raises questions. For instance, how might clinicians address a potential discrepancy between the output of the LLM (the patient's pain score and treatment recommendations) and the patient's testimony? The clinical and ethical concerns are, firstly, that a strong desire for mechanical objectivity in pain management may lead to automation bias, which is the tendency to over-rely on and be unduly confident in algorithmic outputs19; secondly, the clinician might be convinced that any relevant biases have been addressed in the model21; thirdly, any potentially ineffective LLM-generated treatment recommendations might impede the use of more effective options and might also render the patient's clinical status opaque21; fourthly, most patients are not necessarily in a position to verify or challenge the LLM output (or most medical tests) because they are dependent on their clinicians' (and the LLMs') credibility and reliability for their care14; and finally, the clinician may feel compelled to centre the LLM-generated pain score as opposed to the subjective testimony of the patient.11,24 Of course, patient testimony is not the only factor clinicians consider when making patient-centred and clinically appropriate treatment recommendations. Clinicians also weigh factors such as patient medical history, biopsychosocial influences, and patient values and treatment preferences.13 However, clinicians may consider the presumably objective output from the LLM more accurate and reliable, thereby minimizing the patient's credibility and knowledge about their subjective and embodied experiences.2,11,14,24 These issues may be intensified for patients with intersecting disadvantaged identities related to gender, class, sexuality, disability, substance use, or race.3,23,24 Young et al. attempted to create algorithmic fairness by removing demographic factors such as race, ethnicity, and sex from their model. Scholars have cautioned against this approach arguing that it falsely assumes that bias is quantifiable and addressable through algorithmic models. Human biases are often embedded deeply in social and structural contexts, making it difficult to eliminate them solely through technical means.21,26 Furthermore, removing demographic variables in models had limited success previously. For example, some AI tools can predict race or gender incidentally without these features being labeled in the training data but through proxies such as profession or social group associations.8,31 Others have argued that we should take a critical eye to variables such as sex and race in research. For example, biological sex is often conflated with the social construct of gender15 and that labeling race—another complex social construct—as an independent data point obfuscates the harms of structural and interpersonal racism,1 let alone intersectional factors shaping pain experiences and clinical decision-making.20 Addressing systemic biases includes acknowledging and addressing the broader socio-structural factors that contribute to population health inequities in pain management.26,28 The large body of evidence on the social determinants of health demonstrates that antidotes to population health inequities, including pain, exist outside of formal personal healthcare service provision.10,16,17 Using a LLM to solve structural macro-level injustices such as racial, gender-based, and sex-based stratification in pain management at the clinical level should be approached very carefully; it shifts the focus from upstream structural drivers such as public policies, laws, and distribution of health-related risks toward downstream micro-level or individual-level technological solutions to the root causes of population-level inequities.17 I am hopeful that research using AI systems, including LLMs, can make important contributions to the “equity, effectiveness, and efficiency” of pain management.4,18,30 If such research demonstrates evidence of actual patient benefit, it would be invaluable for individual patients as part of comprehensive multimodal pain management. However, the complex structural injustices that Young et al. aim to use LLMs to overcome—racism, sexism, and genderism in pain management—cannot be reduced to computations. A solely technological, algorithmic approach is insufficient to address these urgently important social issues. Conflict of interest statement The author has no conflicts of interest to declare.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,088
score de la tête « metaresearch » (Gemma)0,094
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesÉtudes des sciences et des technologies
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: aucune
GenreSignal candidat: Empirique · Signal consensuel: aucune
Score de désaccord entre enseignants0,995
Score d'incertitude au seuil0,463

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0880,094
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0020,001
Études des sciences et des technologies0,0050,065
Communication savante0,0100,009
Science ouverte0,0030,007
Intégrité de la recherche0,0090,011
Charge utile insuffisante (le modèle a refusé de juger)0,0030,001

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,025
Tête enseignante GPT0,309
Écart entre enseignants0,284 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeSans objet
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations3
Publié2024
Routes d'admission2
Résumé présentoui

Explorer davantage

Même revuePainMême sujetDiversity and Career in MedicineTravaux en français237 207