ChatGPT, Claude, and Gemini all fail a basic real-world test: recommending phones for people with chronic migraines. That’s not speculation — it’s what XDA Developers found when it asked each model the same medically grounded question and checked their answers against known display engineering facts. The results weren’t close calls. They were inconsistent, contradictory, and at times dangerously misleading. And they expose how little guardrails exist between an AI’s confidence and its accuracy, especially on health-adjacent hardware decisions. This isn’t about benchmark scores or camera specs. It’s about whether someone managing migraines can trust an AI to steer them toward a phone that won’t trigger pain. ChatGPT first recommended the Galaxy S26 Ultra. Then later advised against buying it, citing its low PWM dimming rate. Gemini told users to crank brightness to 100% and use third-party dimming apps to avoid flicker, ignoring that full brightness itself is a known migraine trigger, and that screen-overlay apps carry security risks. It also pushed the Honor 200 series, a 2023–2024 lineup, despite no prompt limiting age or budget. Claude avoided naming models altogether. Opting for vague brand-level suggestions, but at least didn’t misstate technical facts. The inconsistency extended beyond phones. When asked whether succulent flower stems can propagate new plants, ChatGPT and Gemini said yes (Gemini without hesitation. ChatGPT hedging with “sometimes”). Claude and Brave Search said no. Correctly. There was no consensus. No shared grounding in botany. Just four different models, trained on different data, making different calls on the same objective fact. None of this is new in theory. Every AI vendor warns users not to treat outputs as truth. But warnings don’t stop people from using ChatGPT or Gemini as quick-reference tools, especially when searching for “chatgpt app” or “gemini ai” on mobile. And those apps are now embedded in search, messaging, and even device settings. Their reach outpaces their reliability. What’s missing isn’t more training data. It’s better signal discipline: knowing when to say “I don’t know,” flagging contested claims, citing primary sources like display white papers or clinical guidelines. Not Reddit threads or outdated press releases. XDA notes that even NotebookLM, built to minimize hallucination, stumbles if fed weak source material. That’s the loop: bad inputs, confident outputs, unverified adoption. For now, there’s no fix baked into the apps themselves. Users still need to cross-check, especially on health, safety, or hardware compatibility. And developers? They’re still shipping AI features faster than they’re auditing their factual boundaries. XDA’s testing ran on current public versions of each model. No rollout dates, firmware updates, or policy changes were announced alongside the findings. The next step isn’t a patch. It’s pressure. From users, regulators, and clinicians, to treat medical-adjacent advice as high-risk output, not just another autocomplete suggestion.