dorsal/arxiv
View SchemaBenchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models
| Authors | Shaonan Liu, Guo Yu, Xiaoling Luo, Shiyi Zheng, Wenting Chen, Jie Liu, Linlin Shen |
|---|---|
| Categories | |
| ArXiv ID | 2601.06750vv1 |
| URL | https://arxiv.org/abs/2601.06750 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Medical Multimodal Large Language Models (Med-MLLMs) require egocentric clinical intent understanding for real-world deployment, yet existing benchmarks fail to evaluate this critical capability. To address these challenges, we introduce MedGaze-Bench, the first benchmark leveraging clinician gaze as a Cognitive Cursor to assess intent understanding across surgery, emergency simulation, and diagnostic interpretation. Our benchmark addresses three fundamental challenges: visual homogeneity of anatomical structures, strict temporal-causal dependencies in clinical workflows, and implicit adherence to safety protocols. We propose a Three-Dimensional Clinical Intent Framework evaluating: (1) Spatial Intent: discriminating precise targets amid visual noise, (2) Temporal Intent: inferring causal rationale through retrospective and prospective reasoning, and (3) Standard Intent: verifying protocol compliance through safety checks. Beyond accuracy metrics, we introduce Trap QA mechanisms to stress-test clinical reliability by penalizing hallucinations and cognitive sycophancy. Experiments reveal current MLLMs struggle with egocentric intent due to over-reliance on global features, leading to fabricated observations and uncritical acceptance of invalid instructions.
{
"annotation_id": "a8f5af5d-362f-4748-b885-b331aa79dae1",
"date_created": "2026-02-17T05:53:07.835000Z",
"date_modified": "2026-02-17T05:53:07.835000Z",
"file_hash": "265e495fe09280058fd699c94e89d2a7fabcb6e59ac90c4cbc3c08fd287b2246",
"private": false,
"record": {
"abstract": "Medical Multimodal Large Language Models (Med-MLLMs) require egocentric clinical intent understanding for real-world deployment, yet existing benchmarks fail to evaluate this critical capability. To address these challenges, we introduce MedGaze-Bench, the first benchmark leveraging clinician gaze as a Cognitive Cursor to assess intent understanding across surgery, emergency simulation, and diagnostic interpretation. Our benchmark addresses three fundamental challenges: visual homogeneity of anatomical structures, strict temporal-causal dependencies in clinical workflows, and implicit adherence to safety protocols. We propose a Three-Dimensional Clinical Intent Framework evaluating: (1) Spatial Intent: discriminating precise targets amid visual noise, (2) Temporal Intent: inferring causal rationale through retrospective and prospective reasoning, and (3) Standard Intent: verifying protocol compliance through safety checks. Beyond accuracy metrics, we introduce Trap QA mechanisms to stress-test clinical reliability by penalizing hallucinations and cognitive sycophancy. Experiments reveal current MLLMs struggle with egocentric intent due to over-reliance on global features, leading to fabricated observations and uncritical acceptance of invalid instructions.",
"arxiv_id": "2601.06750",
"authors": [
"Shaonan Liu",
"Guo Yu",
"Xiaoling Luo",
"Shiyi Zheng",
"Wenting Chen",
"Jie Liu",
"Linlin Shen"
],
"categories": [
"cs.CV",
"cs.AI",
"cs.CL"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models",
"url": "https://arxiv.org/abs/2601.06750",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "d51c33cd-24f4-48b9-bdde-7d02b697e25c",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}