Comparison of ChatGPT and Google Gemini Responses to MS Exercise Questions
In a study assessing AI responses to exercise questions for MS patients, 61.3% of responses from ChatGPT-5.4 were rated high quality, compared to only 24% from Google Gemini 3 Flash (p < 0.001). ChatGPT also outperformed Gemini in accuracy and readability.
- Tags
- #Other
- Most relevant to
- —
Why it matters
These findings highlight that ChatGPT may provide more reliable information for MS patients seeking exercise-related guidance, but they don't imply better health outcomes or real-world applicability of these AI responses.
What this does not prove
This study only evaluated AI-generated responses and did not assess actual patient outcomes or interactions with healthcare professionals.
Next milestone
No next milestone was established from the available source.
Study facts
- Study design
- Observational
- Participants / samples
- 75 · Number of questions evaluated.
- Randomised
- No
- Controlled
- Not reported
- Primary endpoint met
- Not reported
- Relevant MS type
- Not reported
- Publication date
- 2026-09-23
- Evidence reviewed
- Abstract only
- Regulatory approval
- Not reported
- Research areas
- Other
Original sources
- Primary evidenceGenerative AI responses to exercise-related questions asked by patients with multiple sclerosis: An evaluation of quality, accuracy, and readability. ↗DOI: 10.1016/j.msard.2026.107937
Supporting passages (13)
study phaseThis study aims to compare the quality, accuracy, reliability, and readability of responses produced by ChatGPT-5.4 Thinking and Google Gemini 3 Flash regarding frequently asked questions by multiple sclerosis (MS) patients about exercise.
study designA total of 75 questions were evaluated. Expert physiotherapists analysed the responses utilizing the Global Quality Score (GQS), Modified DISCERN (mDISCERN), Likert Accuracy Scale, and the Flesch Reading Ease (FRE).
subjectsThis study aims to compare the quality, accuracy, reliability, and readability of responses produced by ChatGPT-5.4 Thinking and Google Gemini 3 Flash regarding frequently asked questions by multiple sclerosis (MS) patients about exercise.
sample sizeA total of 75 questions were evaluated.
sample size basisMETHOD: A total of 75 questions were evaluated.
randomizedA total of 75 questions were evaluated.
research categoriesOBJECTIVE: This study aims to compare the quality, accuracy, reliability, and readability of responses produced by ChatGPT-5.4 Thinking and Google Gemini 3 Flash regarding frequently asked questions by multiple sclerosis (MS) patients about exercise.
publication datePublication date: 2026-09-23
interventionTitle: Generative AI responses to exercise-related questions asked by patients with multiple sclerosis: An evaluation of quality, accuracy, and readability.
comparatorOBJECTIVE: This study aims to compare the quality, accuracy, reliability, and readability of responses produced by ChatGPT-5.4 Thinking and Google Gemini 3 Flash regarding frequently asked questions by multiple sclerosis (MS) patients about exercise.
primary endpointOBJECTIVE: This study aims to compare the quality, accuracy, reliability, and readability of responses produced by ChatGPT-5.4 Thinking and Google Gemini 3 Flash regarding frequently asked questions by multiple sclerosis (MS) patients about exercise.
findingsRESULTS: While 61.3% of ChatGPT responses were classified as high quality, this rate was 24% for Gemini (p < 0.001).
limitationsTitle: Generative AI responses to exercise-related questions asked by patients with multiple sclerosis: An evaluation of quality, accuracy, and readability.
AI assessment, not yet reviewed by a person · version 1 · Community votes are separate from evidence review.