dorsal/arxiv
View SchemaART: Action-based Reasoning Task Benchmarking for Medical AI Agents
| Authors | Ananya Mantravadi, Shivali Dalmia, Abhishek Mukherji |
|---|---|
| Categories | |
| ArXiv ID | 2601.08988vv1 |
| URL | https://arxiv.org/abs/2601.08988 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Reliable clinical decision support requires medical AI agents capable of safe, multi-step reasoning over structured electronic health records (EHRs). While large language models (LLMs) show promise in healthcare, existing benchmarks inadequately assess performance on action-based tasks involving threshold evaluation, temporal aggregation, and conditional logic. We introduce ART, an Action-based Reasoning clinical Task benchmark for medical AI agents, which mines real-world EHR data to create challenging tasks targeting known reasoning weaknesses. Through analysis of existing benchmarks, we identify three dominant error categories: retrieval failures, aggregation errors, and conditional logic misjudgments. Our four-stage pipeline -- scenario identification, task generation, quality audit, and evaluation -- produces diverse, clinically validated tasks grounded in real patient data. Evaluating GPT-4o-mini and Claude 3.5 Sonnet on 600 tasks shows near-perfect retrieval after prompt refinement, but substantial gaps in aggregation (28--64%) and threshold reasoning (32--38%). By exposing failure modes in action-oriented EHR reasoning, ART advances toward more reliable clinical agents, an essential step for AI systems that reduce cognitive load and administrative burden, supporting workforce capacity in high-demand care settings
{
"annotation_id": "9f2e94cd-0639-4c2e-9bb3-f012453b3736",
"date_created": "2026-02-17T05:53:20.022000Z",
"date_modified": "2026-02-17T05:53:20.022000Z",
"file_hash": "42c3d5cab6da3e24100c25a30b47e7796d62c49aad08698bdd62b9c4c4190483",
"private": false,
"record": {
"abstract": "Reliable clinical decision support requires medical AI agents capable of safe, multi-step reasoning over structured electronic health records (EHRs). While large language models (LLMs) show promise in healthcare, existing benchmarks inadequately assess performance on action-based tasks involving threshold evaluation, temporal aggregation, and conditional logic. We introduce ART, an Action-based Reasoning clinical Task benchmark for medical AI agents, which mines real-world EHR data to create challenging tasks targeting known reasoning weaknesses. Through analysis of existing benchmarks, we identify three dominant error categories: retrieval failures, aggregation errors, and conditional logic misjudgments. Our four-stage pipeline -- scenario identification, task generation, quality audit, and evaluation -- produces diverse, clinically validated tasks grounded in real patient data. Evaluating GPT-4o-mini and Claude 3.5 Sonnet on 600 tasks shows near-perfect retrieval after prompt refinement, but substantial gaps in aggregation (28--64%) and threshold reasoning (32--38%). By exposing failure modes in action-oriented EHR reasoning, ART advances toward more reliable clinical agents, an essential step for AI systems that reduce cognitive load and administrative burden, supporting workforce capacity in high-demand care settings",
"arxiv_id": "2601.08988",
"authors": [
"Ananya Mantravadi",
"Shivali Dalmia",
"Abhishek Mukherji"
],
"categories": [
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "ART: Action-based Reasoning Task Benchmarking for Medical AI Agents",
"url": "https://arxiv.org/abs/2601.08988",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "8d28a118-0a7f-4352-aae6-22fb894e49fa",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}