Ask ChatGPT about a weird symptom and you’re in enormous company. About 32% of U.S. adults have used an AI chatbot for health information in the past year, according to KFF’s 2026 tracking poll, roughly the same share who turn to social media for the same reason. OpenAI’s own numbers are larger still: more than 40 million people ask ChatGPT a healthcare question every day, and healthcare now makes up over 5% of all ChatGPT messages worldwide. That same OpenAI report found three in five U.S. adults had used AI for a health question in the past three months, most often to explore symptoms (55%) or make sense of medical terms (48%).
Scott Shimp, VP of Solutions at Data Society, doesn’t think that habit itself is the problem. What worries him is what happens right after.
“A classic case here is that you might ask some AI tool about a medical issue,” he says. “You might say, ‘Hey ChatGPT, or Claude, or whatever, what do you think of this? What should I do? What’s wrong with me? What should I do next?’ In that situation, you should ask yourself: is there some kind of expert or institution that would be the natural go-to for that question? And even if you’ve convinced yourself that there’s urgency, and for whatever reason you don’t want to go talk to, in this case, a doctor or a medical professional, you can still reframe it a little bit and think: okay, I, myself, as an outside person here looking at the situation, would I trust myself having made that decision only on the basis of what I’m hearing from the AI tool? And then you can try, to the extent you can, to objectively look at the situation and think about: is this really enough for me to call it case closed and go with the AI on this one?”
The same KFF poll found that 58% of people who asked AI about a mental health question, and 42% who asked about a physical health question, never followed up with a doctor afterward. They asked, got an answer, and stopped there, which is exactly the gap Shimp is describing.
Stopping there is riskiest exactly when it feels safest. A UCLA-led study published through BMJ audited five popular chatbots (Gemini, DeepSeek, Meta AI, ChatGPT, and Grok) on real medical questions and found 49.6% of responses were problematic: 30% somewhat, 19.6% highly. References were weak across the board, with a median citation-completeness score of just 40%, and every chatbot tested fabricated at least some citations outright. The researchers found the answers, right and wrong alike, were delivered “with confidence and certainty, with few caveats or disclaimers,” the exact tone that makes case closed feel reasonable even when it isn’t.
Shimp’s test for catching that moment has two parts, and neither requires distrusting AI outright. First, ask whether there’s an expert or institution that’s the natural go-to for this specific question. If the honest answer is yes, that’s worth noticing before you skip it. Second, for the cases where you’ve decided not to make that call anyway, imagine yourself as an outside observer. Would you trust someone else who made this decision based only on what an AI tool told them? If the answer is no, the AI’s answer wasn’t wrong to ask for. It just wasn’t enough to stop on.
Shimp uses health as his example because it’s the clearest case, but the test travels. The same two questions apply to legal, financial, or any other question where an obvious human expert exists: is there a natural place to take this, and would you stand behind a decision made only on the AI’s word?
Try Shimp’s two-question test on the next AI answer you’re tempted to treat as final, medical or otherwise: is there a natural expert for this question, and would you trust someone else who stopped where you’re about to stop? For teams building that kind of judgment across an organization, Data Society’s AI upskilling programs teach exactly this kind of decision-making. And for organizations that need a framework for when AI output gets escalated to a person, AI Advisory is where that conversation starts.
Frequently Asked Questions
Asking is a reasonable first move, and Scott Shimp does not treat the habit itself as the problem. The risk starts the moment you accept the answer as final and skip the expert you would otherwise have called.
Often they do not, according to that same KFF poll. Among people who asked AI a mental health question, 58% never followed up with a doctor afterward, along with 42% of those who asked about physical health.
A UCLA-led audit published in BMJ Open rated 49.6% of chatbot responses as problematic, split between 30% somewhat problematic and 19.6% highly problematic. Reference quality was weak across all five chatbots tested, at a median citation-completeness score of 40%, and every one of them fabricated citations outright.
First, ask whether an expert or institution is the natural place to take this specific question. Second, ask whether you would trust another person who made the same decision based only on what an AI told them.

