dorsal/arxiv
View SchemaGROKE: Vision-Free Navigation Instruction Evaluation via Graph Reasoning on OpenStreetMap
| Authors | Farzad Shami, Subhrasankha Dey, Nico Van de Weghe, Henrikki Tenkanen |
|---|---|
| Categories | |
| ArXiv ID | 2601.07375vv1 |
| URL | https://arxiv.org/abs/2601.07375 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
The evaluation of navigation instructions remains a persistent challenge in Vision-and-Language Navigation (VLN) research. Traditional reference-based metrics such as BLEU and ROUGE fail to capture the functional utility of spatial directives, specifically whether an instruction successfully guides a navigator to the intended destination. Although existing VLN agents could serve as evaluators, their reliance on high-fidelity visual simulators introduces licensing constraints and computational costs, and perception errors further confound linguistic quality assessment. This paper introduces GROKE(Graph-based Reasoning over OSM Knowledge for instruction Evaluation), a vision-free training-free hierarchical LLM-based framework for evaluating navigation instructions using OpenStreetMap data. Through systematic ablation studies, we demonstrate that structured JSON and textual formats for spatial information substantially outperform grid-based and visual graph representations. Our hierarchical architecture combines sub-instruction planning with topological graph navigation, reducing navigation error by 68.5% compared to heuristic and sampling baselines on the Map2Seq dataset. The agent's execution success, trajectory fidelity, and decision patterns serve as proxy metrics for functional navigability given OSM-visible landmarks and topology, establishing a scalable and interpretable evaluation paradigm without visual dependencies. Code and data are available at https://anonymous.4open.science/r/groke.
{
"annotation_id": "87221de6-85ca-44e6-bdfa-e1210b49ef27",
"date_created": "2026-02-17T05:53:12.387000Z",
"date_modified": "2026-02-17T05:53:12.387000Z",
"file_hash": "3bc69e89a328bdaa770ef726e1df85ec6b4ddbdeed81bc1eedec828512592687",
"private": false,
"record": {
"abstract": "The evaluation of navigation instructions remains a persistent challenge in Vision-and-Language Navigation (VLN) research. Traditional reference-based metrics such as BLEU and ROUGE fail to capture the functional utility of spatial directives, specifically whether an instruction successfully guides a navigator to the intended destination. Although existing VLN agents could serve as evaluators, their reliance on high-fidelity visual simulators introduces licensing constraints and computational costs, and perception errors further confound linguistic quality assessment. This paper introduces GROKE(Graph-based Reasoning over OSM Knowledge for instruction Evaluation), a vision-free training-free hierarchical LLM-based framework for evaluating navigation instructions using OpenStreetMap data. Through systematic ablation studies, we demonstrate that structured JSON and textual formats for spatial information substantially outperform grid-based and visual graph representations. Our hierarchical architecture combines sub-instruction planning with topological graph navigation, reducing navigation error by 68.5% compared to heuristic and sampling baselines on the Map2Seq dataset. The agent\u0027s execution success, trajectory fidelity, and decision patterns serve as proxy metrics for functional navigability given OSM-visible landmarks and topology, establishing a scalable and interpretable evaluation paradigm without visual dependencies. Code and data are available at https://anonymous.4open.science/r/groke.",
"arxiv_id": "2601.07375",
"authors": [
"Farzad Shami",
"Subhrasankha Dey",
"Nico Van de Weghe",
"Henrikki Tenkanen"
],
"categories": [
"cs.CL"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "GROKE: Vision-Free Navigation Instruction Evaluation via Graph Reasoning on OpenStreetMap",
"url": "https://arxiv.org/abs/2601.07375",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "db4e0d64-2de9-47c7-9aad-e3fb44eec177",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}