Page 14 - hypertension_newsletter
P. 14

REFLECTIONS
                                                                                                                   Hypertension
     Hypertension Global Newsletter #10 2026


     Three blinded expert reviewers then rated the four responses on accuracy and safety, and attempted to identify the source
     of each response (human vs. AI).                                                                              Hypertension

     The findings revealed that while GPT-4 outperformed the other models, it still fell short of human expert standards. Specifically,
     GPT-4 achieved an overall accurate response rate of 83% and a safety rate of 86%, compared to the expert’s 92% accuracy
     and 93% safety. The other models performed significantly worse; Gemini achieved 64% accuracy, and MedLM (despite being
     specifically designed for the healthcare industry) scored 35% for accuracy and 39% for safety. Furthermore, less than half
     (46%) of GPT-4’s responses were deemed fully guideline-concordant, compared to the expert’s 68%.


                             Percentage of accurate responses according to the source of response















































     Reviewers struggled to differentiate between human and AI-generated responses. They incorrectly identified 25% of the
     expert’s answers as coming from a chatbot, and they could only correctly identify GPT-4 as an AI 46% of the time. Ultimately,
     the study highlights a key challenge: LLMs can successfully replicate human-like reasoning and authoritative tone, which
     heightens the risk of clinicians overly trusting content that may secretly diverge from safe, evidence-based practices.


     While advanced LLMs like GPT-4 show immense potential to eventually be integrated into electronic health records to reduce
     therapeutic inertia and assist primary care clinicians at the point of care, current versions remain inferior to human clinical
     judgment. Therefore, human-in-the-loop supervision,
     rigorous benchmarking, and domain-specific fine-tuning
     remain essential prerequisites before LLMs can be safely             CLICK HERE
     deployed for real-world clinical decision-making.                    FOR THE LINK TO FULL ARTICLE



          TABLE OF CONTENTS
   9   10   11   12   13   14   15   16   17   18   19