NavAble: A Large-Scale Dataset and Synthetic Data Generation Pipeline for Blind Navigation

NeurIPS 2026 ED Track

Hochul Hwang†, Jahir Sadik Monon†, Soowan Yang, Shiven Umeshbhai Patel, Anh Nhat Hong Nguyen, Kien Nguyen, Keshav Garg, Dylan Gage, Duretti Hordofaa, Anshu Anjna, Eshed Ohn-Bar, Donghyun Kim

† Equal contribution

Perceiving the objects that make navigation accessible.
NavAble brings together real-world data and controllable synthetic generation for blind and low-vision navigation.

11
accessibility object classes
452K+
synthetic frames
37
simulation environments
500
released 3D assets

Abstract

Reliable recognition of accessibility-critical objects, such as accessible pedestrian signals, door-activation buttons, and handrails, is essential for technologies that support independent mobility for blind and low-vision (BLV) people. Yet these objects are often missing or underrepresented in existing vision datasets, and available data offer limited diversity, inconsistent annotations, and viewpoints that do not match assistive devices.

We introduce NavAble, a large-scale dataset and synthetic data generation pipeline for BLV navigation. NavAble combines curated open-source images with newly collected, densely annotated real-world data across 11 accessibility-critical categories. Its generator transforms web images into simulation-ready 3D assets and renders them in NVIDIA Isaac Sim under varied environments, lighting, and camera viewpoints, producing RGB images with rich ground-truth annotations.

We evaluate accessibility-object segmentation across multiple model architectures and study how synthetic supervision supports real-world generalization. Evaluation on a quadruped guide dog robot and a cane-shaped mobility-assistive device further examines transfer to the egocentric viewpoints encountered during navigation.

Key Contributions

The NavAble Dataset

Three complementary data sources expand coverage of the objects that support everyday mobility decisions.

ComponentImagesRole
Curated real-world data36,466Existing sources remapped to a shared class taxonomy.
Newly collected real-world data5,581Dense annotations from indoor and outdoor navigation settings.
Synthetic data452,704Controllable appearance, lighting, environments, and viewpoints.

Counts reflect the public dataset release. The collected real-world test split contains 1,482 images.

APS buttonBus stopBus stop signCrosswalkDoor buttonElevatorElevator buttonEscalatorHandrailPedestrian signalTurnstile

APS: accessible pedestrian signal. Turnstile is synthetic-only in the current release; real-world evaluation uses the 10 shared classes.

Examples of inconsistent annotations and noisy open-source images alongside curated and newly collected images with dense, consistent masks.
Careful curation and targeted collection address noisy labels, missing classes, and inconsistent masks. Select a figure to view it at full size.

From Images to Synthetic Training Data

Specify an object, place an asset, and record a camera trajectory. Replay that trajectory across object variants, lighting conditions, and environments.

NavAble generator: web images are filtered by a vision-language model, grounded and masked, reconstructed into 3D objects, then rendered with configurable asset positions and camera trajectories across day, sunset, and night environments.
The NavAble Generator combines foundation-model-based asset generation with configurable rendering in NVIDIA Isaac Sim.
  1. Find and filter. Retrieve candidate object images and filter unsuitable examples with a vision-language model.
  2. Reconstruct assets. Localize and segment objects, then reconstruct textured 3D assets with SAM 3D Objects.
  3. Configure a scene. Set the asset placement and record a reusable camera trajectory.
  4. Render variations. Generate RGB, semantic segmentation, depth, surface normals, and 2D/3D bounding boxes.

Video

Lighting variation, annotation modalities, and a nine-object montage from the synthetic data generator.

The video first shows a pedestrian-signal button under changing illumination, then paired renderings and annotations, and finally nine examples of accessibility-related infrastructure.

Segmentation Evaluation

We benchmark SAN, Mask2Former, SegFormer, DeepLabv3+, and EncNet on a held-out real-world test set. Four training configurations separate the effects of newly collected data, curated data, and synthetic supervision.

RealReal + CuratedReal + SyntheticReal + Curated + Synthetic

The results show architecture-dependent effects: synthetic augmentation improves the evaluated transformer and mask-classification models, while the CNN baselines do not improve over their real-only baselines. This comparison helps identify where synthetic diversity is useful and where domain differences remain challenging.

Transfer to Assistive Robots

From dataset evaluation to the viewpoints of a quadruped guide dog robot and a cane-shaped mobility-assistive device.

Cane-shaped and quadruped assistive platforms, with ground-truth masks and SegFormer predictions comparing training on real data only against real plus synthetic data in navigation scenes.
Qualitative evaluation on egocentric footage from two mobility-assistive platforms. Each comparison pairs real-only training with real plus synthetic supervision.

Using SegFormer on both platforms, we evaluate whether improvements transfer to previously unseen device viewpoints. Synthetic supervision improves mean IoU and precision on both embodiment splits, with benefits for navigation-relevant objects such as pedestrian signals, APS buttons, and handrails.

These experiments evaluate perception transfer. Reliable closed-loop navigation remains an open direction, alongside broader class coverage and more geographically diverse real-world data.