dorsal/arxiv
View SchemaGI-Bench: A Panoramic Benchmark Revealing the Knowledge-Experience Dissociation of Multimodal Large Language Models in Gastrointestinal Endoscopy Against Clinical Standards
| Authors | Yan Zhu, Te Luo, Pei-Yao Fu, Zhen Zhang, Zi-Long Wang, Yi-Fan Qu, Zi-Han Geng, Jia-Qi Xu, Lu Yao, Li-Yun Ma, Wei Su, Wei-Feng Chen, Quan-Lin Li, Shuo Wang, Ping-Hong Zhou |
|---|---|
| Categories | |
| ArXiv ID | 2601.08183vv2 |
| URL | https://arxiv.org/abs/2601.08183 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Multimodal Large Language Models (MLLMs) show promise in gastroenterology, yet their performance against comprehensive clinical workflows and human benchmarks remains unverified. To systematically evaluate state-of-the-art MLLMs across a panoramic gastrointestinal endoscopy workflow and determine their clinical utility compared with human endoscopists. We constructed GI-Bench, a benchmark encompassing 20 fine-grained lesion categories. Twelve MLLMs were evaluated across a five-stage clinical workflow: anatomical localization, lesion identification, diagnosis, findings description, and management. Model performance was benchmarked against three junior endoscopists and three residency trainees using Macro-F1, mean Intersection-over-Union (mIoU), and multi-dimensional Likert scale. Gemini-3-Pro achieved state-of-the-art performance. In diagnostic reasoning, top-tier models (Macro-F1 0.641) outperformed trainees (0.492) and rivaled junior endoscopists (0.727; p>0.05). However, a critical "spatial grounding bottleneck" persisted; human lesion localization (mIoU >0.506) significantly outperformed the best model (0.345; p<0.05). Furthermore, qualitative analysis revealed a "fluency-accuracy paradox": models generated reports with superior linguistic readability compared with humans (p<0.05) but exhibited significantly lower factual correctness (p<0.05) due to "over-interpretation" and hallucination of visual features. GI-Bench maintains a dynamic leaderboard that tracks the evolving performance of MLLMs in clinical endoscopy. The current rankings and benchmark results are available at https://roterdl.github.io/GIBench/.
{
"annotation_id": "8bfc8bde-fa61-47e7-bc04-84ede0265605",
"date_created": "2026-02-17T05:53:15.938000Z",
"date_modified": "2026-02-17T05:53:15.938000Z",
"file_hash": "cd5c6d1fda72d9337838dc022686df4c44f904f91a8b5780adfb1433ac865a85",
"private": false,
"record": {
"abstract": "Multimodal Large Language Models (MLLMs) show promise in gastroenterology, yet their performance against comprehensive clinical workflows and human benchmarks remains unverified. To systematically evaluate state-of-the-art MLLMs across a panoramic gastrointestinal endoscopy workflow and determine their clinical utility compared with human endoscopists. We constructed GI-Bench, a benchmark encompassing 20 fine-grained lesion categories. Twelve MLLMs were evaluated across a five-stage clinical workflow: anatomical localization, lesion identification, diagnosis, findings description, and management. Model performance was benchmarked against three junior endoscopists and three residency trainees using Macro-F1, mean Intersection-over-Union (mIoU), and multi-dimensional Likert scale. Gemini-3-Pro achieved state-of-the-art performance. In diagnostic reasoning, top-tier models (Macro-F1 0.641) outperformed trainees (0.492) and rivaled junior endoscopists (0.727; p\u003e0.05). However, a critical \"spatial grounding bottleneck\" persisted; human lesion localization (mIoU \u003e0.506) significantly outperformed the best model (0.345; p\u003c0.05). Furthermore, qualitative analysis revealed a \"fluency-accuracy paradox\": models generated reports with superior linguistic readability compared with humans (p\u003c0.05) but exhibited significantly lower factual correctness (p\u003c0.05) due to \"over-interpretation\" and hallucination of visual features. GI-Bench maintains a dynamic leaderboard that tracks the evolving performance of MLLMs in clinical endoscopy. The current rankings and benchmark results are available at https://roterdl.github.io/GIBench/.",
"arxiv_id": "2601.08183",
"authors": [
"Yan Zhu",
"Te Luo",
"Pei-Yao Fu",
"Zhen Zhang",
"Zi-Long Wang",
"Yi-Fan Qu",
"Zi-Han Geng",
"Jia-Qi Xu",
"Lu Yao",
"Li-Yun Ma",
"Wei Su",
"Wei-Feng Chen",
"Quan-Lin Li",
"Shuo Wang",
"Ping-Hong Zhou"
],
"categories": [
"cs.CV",
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "GI-Bench: A Panoramic Benchmark Revealing the Knowledge-Experience Dissociation of Multimodal Large Language Models in Gastrointestinal Endoscopy Against Clinical Standards",
"url": "https://arxiv.org/abs/2601.08183",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "3a83eddc-d48c-4baf-825e-1f0cd63fdee7",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}