dorsal/arxiv
View SchemaData Scaling for Navigation in Unknown Environments
| Authors | Lauri Suomela, Naoki Takahata, Sasanka Kuruppu Arachchige, Harry Edelman, Joni-Kristian Kämäräinen |
|---|---|
| Categories | |
| ArXiv ID | 2601.09444vv1 |
| URL | https://arxiv.org/abs/2601.09444 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Generalization of imitation-learned navigation policies to environments unseen in training remains a major challenge. We address this by conducting the first large-scale study of how data quantity and data diversity affect real-world generalization in end-to-end, map-free visual navigation. Using a curated 4,565-hour crowd-sourced dataset collected across 161 locations in 35 countries, we train policies for point goal navigation and evaluate their closed-loop control performance on sidewalk robots operating in four countries, covering 125 km of autonomous driving. Our results show that large-scale training data enables zero-shot navigation in unknown environments, approaching the performance of policies trained with environment-specific demonstrations. Critically, we find that data diversity is far more important than data quantity. Doubling the number of geographical locations in a training set decreases navigation errors by ~15%, while performance benefit from adding data from existing locations saturates with very little data. We also observe that, with noisy crowd-sourced data, simple regression-based models outperform generative and sequence-based architectures. We release our policies, evaluation setup and example videos on the project page.
{
"annotation_id": "8a60c7b8-23d8-475e-8ddc-9cfdff1e9f92",
"date_created": "2026-02-17T05:53:20.331000Z",
"date_modified": "2026-02-17T05:53:20.331000Z",
"file_hash": "0a15c0012007f68adefe0cca31fb12b305299ba881945c921dd55ece09bc553b",
"private": false,
"record": {
"abstract": "Generalization of imitation-learned navigation policies to environments unseen in training remains a major challenge. We address this by conducting the first large-scale study of how data quantity and data diversity affect real-world generalization in end-to-end, map-free visual navigation. Using a curated 4,565-hour crowd-sourced dataset collected across 161 locations in 35 countries, we train policies for point goal navigation and evaluate their closed-loop control performance on sidewalk robots operating in four countries, covering 125 km of autonomous driving.\n Our results show that large-scale training data enables zero-shot navigation in unknown environments, approaching the performance of policies trained with environment-specific demonstrations. Critically, we find that data diversity is far more important than data quantity. Doubling the number of geographical locations in a training set decreases navigation errors by ~15%, while performance benefit from adding data from existing locations saturates with very little data. We also observe that, with noisy crowd-sourced data, simple regression-based models outperform generative and sequence-based architectures. We release our policies, evaluation setup and example videos on the project page.",
"arxiv_id": "2601.09444",
"authors": [
"Lauri Suomela",
"Naoki Takahata",
"Sasanka Kuruppu Arachchige",
"Harry Edelman",
"Joni-Kristian K\u00e4m\u00e4r\u00e4inen"
],
"categories": [
"cs.RO"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Data Scaling for Navigation in Unknown Environments",
"url": "https://arxiv.org/abs/2601.09444",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "b7215f84-1a6c-47cc-970c-31c8008a2769",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}