- LLMs show promise in hepatology for text and image tasks including radiology, clinical decision support, and patient education.
- Accuracy and reliability are highly variable across tasks, with hallucinations and dependence on training data quality limiting clinical safety.
- Further research is needed on training methods, data quality, ethics, and real world validation against standardised benchmarks before clinical integration.
Hepatol Forum. 2026 Apr 7;7(3):227-235. doi: 10.14744/hf.2026.65328. eCollection 2026.
ABSTRACT
BACKGROUND AND AIM: The rapid advancement of generative artificial intelligence (AI), particularly large language models (LLMs), has opened new frontiers in healthcare, with emerging implications for hepatology. This systematic review synthesizes the current state of research on the application of LLMs in hepatology, focusing on their capabilities in real-world clinical settings, limitations, and future directions.
MATERIALS AND METHODS: Electronic databases, including MEDLINE, EMBASE, and OVID as a search platform, were used to identify eligible studies from inception to January 2025. Eligible studies investigated the clinical utility and performance of LLMs in hepatology, with a clear comparison to a defined ground truth. Key findings were extracted and synthesized narratively. The ROBINS-I tool was used to assess the risk of bias in each study.
RESULTS: Twenty-one studies were included in this review. Our analysis reveals that LLMs demonstrate promising capabilities in processing textual and visual data related to various liver diseases, including hepatocellular carcinoma, cirrhosis, and non-alcoholic fatty liver disease. LLMs effectively assisted with radiological image interpretation, provided clinical decision support, and generated patient education materials. However, the accuracy of these models was highly variable, depending on the specific task and the complexity of the clinical scenario. Limitations, such as the generation of inaccurate or misleading information (“hallucinations”), dependence on training data quality, and ethical considerations, were identified across multiple studies.
CONCLUSION: Generative AI demonstrates feasibility across various hepatology applications, but study heterogeneity and significant challenges remain regarding accuracy, reliability, and safety. Future integration necessitates further research into training methods, data quality, ethical considerations, and real-world validation against standardized benchmarks.
PMID:42602012 | PMC:PMC13473248 | DOI:10.14744/hf.2026.65328
Share Evidence Blueprint
Save to Google Notes

Search Google Scholar
Save as PDF
⭐ My Revision List

