Page 14 - hypertension_newsletter
P. 14
REFLECTIONS
Hypertension
Hypertension Global Newsletter #10 2026
Three blinded expert reviewers then rated the four responses on accuracy and safety, and attempted to identify the source
of each response (human vs. AI). Hypertension
The findings revealed that while GPT-4 outperformed the other models, it still fell short of human expert standards. Specifically,
GPT-4 achieved an overall accurate response rate of 83% and a safety rate of 86%, compared to the expert’s 92% accuracy
and 93% safety. The other models performed significantly worse; Gemini achieved 64% accuracy, and MedLM (despite being
specifically designed for the healthcare industry) scored 35% for accuracy and 39% for safety. Furthermore, less than half
(46%) of GPT-4’s responses were deemed fully guideline-concordant, compared to the expert’s 68%.
Percentage of accurate responses according to the source of response
Reviewers struggled to differentiate between human and AI-generated responses. They incorrectly identified 25% of the
expert’s answers as coming from a chatbot, and they could only correctly identify GPT-4 as an AI 46% of the time. Ultimately,
the study highlights a key challenge: LLMs can successfully replicate human-like reasoning and authoritative tone, which
heightens the risk of clinicians overly trusting content that may secretly diverge from safe, evidence-based practices.
While advanced LLMs like GPT-4 show immense potential to eventually be integrated into electronic health records to reduce
therapeutic inertia and assist primary care clinicians at the point of care, current versions remain inferior to human clinical
judgment. Therefore, human-in-the-loop supervision,
rigorous benchmarking, and domain-specific fine-tuning
remain essential prerequisites before LLMs can be safely CLICK HERE
deployed for real-world clinical decision-making. FOR THE LINK TO FULL ARTICLE
TABLE OF CONTENTS

