Generative Artificial Intelligence in Hip and Knee Arthroplasty: A Systematic Review of Emerging Clinical Applications in Patient Communication and Education, Documentation, and Decision Support
Synopsis
This systematic review searched PubMed and Embase (July 9, 2025) and included 23 studies to assess generative AI, mainly ChatGPT 3.5/4, in total hip and knee arthroplasty across patient communication and education (n=19), clinical documentation (n=2), and clinical decision support (n=2): blinded ratings found FAQ responses comparable to surgeon-written answers in accuracy, clarity, and completeness with better readability; consent documents showed better readability and completeness than surgeon versions; operative-report extraction reached 97.5%–100% agreement; decision support showed higher accuracy for surgical candidacy but low specificity for outcome prediction, alongside fabricated citations and limited patient trust.
Bar chart depicting yearly PubMed-indexed publications using the search string: (“ChatGPT” OR “large language model” OR “generative AI”) AND (“hip arthroplasty” OR “knee arthroplasty” OR “THA OR TKA”). There were 4 articles in 2023, 22 in 2024, and 21 in 2025 (as of July). THA = total hip arthroplasty and TKA = total knee arthroplasty.
PubMedInterpretation
In patient communication and education (19 studies), blinded clinician ratings showed ChatGPT-generated responses to frequently asked questions were comparable to surgeon-written responses in accuracy, clarity, and completeness, with some studies reporting higher empathy, accuracy, and overall quality than 3 of 5 attending surgeons ("4.4/5 vs. <4; p < 0.001"), and educational materials were rewritten from a grade 12 to a grade 6 reading level. First consolidation of THA/TKA generative AI studies into three clinical domains; previously "no systematic review has evaluated generative AI use in THA/TKA specifically." Dominated by blinded Likert ratings, DISCERN scores, and readability scores; heterogeneity across studies was high, synthesis was narrative, and no formal risk-of-bias assessment was performed.
In clinical documentation (2 studies), ChatGPT-4 extracted fixation, technology, and approach information from 240 THA reports with 100%, 98.9%, and 97.5% agreement, with 87% of concise rationales consistent with the text; AI-generated risk-benefit-alternative consent documents outperformed surgeon versions (reading grade 12.6 vs. 16.8; completeness and accuracy 2.4/3 vs. 1.8/3). Combines operative-report information capture and informed-consent drafting as documentation tasks in the arthroplasty setting, suggesting reduced administrative workload. Small number of studies (n=2), but specific agreement rates and score comparisons are reported.
In clinical decision support (2 studies), ChatGPT-3.5 agreed with the consensus of 73 surgeons in 84% of 32 unicompartmental versus total knee replacement cases (95% CI 0.67–0.94), with 91% sensitivity, 70% specificity, and Cohen's κ=0.63; in an 80-patient TKA outcome sample it detected improvement with 97.4% sensitivity but only 33.3% specificity and 65% accuracy, versus surgeons' 90%, 63%, and 76%. Shows a split picture in this domain: candidacy judgments align with expert consensus, while inferring outcomes from unstructured reports has markedly low specificity; overall evidence is "mixed." Samples of 32 and 80 patients with confidence intervals and κ reported, but no formal risk-of-bias assessment.
On trust and citation reliability, patients preferred ChatGPT responses (53.8% vs. 34.0% for nurses; 69.4% in another study), yet 92.1% and 79.0% were unsure about trusting AI in this context; a citation check found only 35.8% of 109 references valid and 64.2% erroneous, while another study reported 37% fabricated or unverifiable references. Frames patient trust and citation fabrication as key watch points, distinguishing patient education, where physicians review content, from decision support, where sources matter more. Based on survey preference counts and reference verification; descriptive results without prospective validation.
Perspective
The review addresses outpatient and perioperative hip and knee arthroplasty settings: informing answers to common patient questions, rewriting educational materials, drafting informed-consent documents, and extracting operative-report information; its audience is arthroplasty surgeons, nursing teams, and hospital informatics departments. Findings apply to an English-language literature base centered on ChatGPT 3.5/4 and rest on human scoring and rating instruments rather than prospective deployment in real clinical workflows.
Models improve quickly, and the text notes that "the results indicate at least 1-year-old performance"; only 3 of 23 articles analyzed models other than ChatGPT, limiting cross-model comparison; studies differ in methodology, sample size, and outcome measures, and no formal risk-of-bias or certainty-of-evidence assessment was performed; patient trust and HIPAA compliance still await prospective validation. The full text of this review is loaded, but the supplemental material (Appendices 1 and 2) and Figure 2 are not included, so screening flow and per-study data extraction cannot be verified here.
