Skin disease AI study shows explainable methods can mislead non-experts

Skin disease AI assistance lands differently depending on who is using it. A study published today in Nature Medicine from MIT, Stanford, and Columbia finds that explanations designed to help users interpret an AI prediction can backfire on the very people they are meant to serve, while giving clinicians little they did not already do better on their own.

The paper, led by Orson Xu of Columbia with Marzyeh Ghassemi of MIT and Roxana Daneshjou of Stanford as co-authors, divided users into two groups. Non-experts decided whether a skin mole image looked cancerous. Primary care clinicians did a harder task: a differential diagnosis across dermatological disease. Both groups saw medical images with an AI prediction, then were given one of four explanation styles layered on top. The conditions included a confidence score alone, similar images that reinforced the prediction, a heat map highlighting important regions, and a plain-language rationale generated by an LLM.

For non-experts, every explanation style improved accuracy, but the underlying mechanism was deference rather than improved reasoning. The researchers found that non-experts accepted LLM explanations whether those explanations were right or wrong, and rated explanations as more persuasive when they sounded vague or generic. The LLM treatment was the largest source of what the paper calls a deference effect, and non-experts reported higher confidence in wrong answers when given an LLM rationale than when given only the model's prediction.

Clinicians showed a pattern that ran the other way. They were not thrown off by incorrect AI explanations, and the prediction-only condition produced the best performance for them. Of the four explanation styles, the LLM treatment boosted clinician accuracy the least, which the paper attributes to clinicians arriving with a diagnosis hypothesis already in mind that they tested the AI against. A plausible but wrong rationale could not pull a clinician's answer the way it could pull a non-expert's.

This expertise asymmetry, where the same intervention produces opposite effects across users, is the editorial core of the study. The non-experts most likely to benefit from AI help are also the most likely to anchor on a confident, fluent rationale. The authors flag this directly by pointing out that users who deferred most to the AI were also the worst performers when stripped of AI assistance, suggesting that the explanation was substituting for reasoning rather than augmenting it.

The team also tested whether timing matters. When an explanation arrived before the user formed their own judgment, deference was stronger than when the explanation followed the user's initial call. That ordering effect matters for any consumer tool that presents AI reasoning up front, and the researchers propose a different default: ask users to commit to a hypothesis first, then offer AI suggestions as alternative diagnoses worth considering rather than as primary answers.

A separate finding complicates the design choice further. The researchers found that AI systems outperformed humans on subtle disease presentations, while humans outperformed the AI on atypical features or noisy images. The implication is that the worst cases for the AI, atypical or noisy presentations, may also be the cases where non-expert deference is most dangerous, since the user has no training to fall back on when the model's prediction is built on weak or unrepresentative inputs.

There is also a fairness dimension. The paper reports that a fairness-constrained model improved accuracy and reduced diagnostic disparities across skin tones. That suggests the deference problem can be partly addressed at the model level, but it does not address the downstream issue of users following a wrong explanation, which is a UX problem rather than a model problem.

For direct-to-consumer skin disease tools and clinical decision-support products deployed without a clinician in the loop, the practical concern is whether the study's effects transfer beyond the dermatology setting tested. The Nature Medicine study finds the relationship between explanations and accuracy is expertise-dependent: a confident, well-written rationale can be a liability as easily as an asset, and the LLM treatment was the worst for non-experts. The paper does not test whether forced-hypothesis workflows close the deference gap in production or under repeated use, so the recommended design remains a proposal rather than a validated intervention. The downstream measurement problem, detecting deference directly in task telemetry, is what productizing these findings depends on.

Subscribe to AI Enthusiast Log

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe