Toggle Main Menu Toggle Search

Open Access padlockePrints

The Newcastle University research output collection, currently available on ePrints, will shortly be moving to a new open repository platform, Figshare. To prepare for the data migration we have paused adding new content to ePrints, and will resume once the new repository is launched. During this time you will continue to have access to ePrints (but no new content will appear). We will share updates here when available.

CoIN: Interactive Navigation With Counterfactual Reasoning via Vision-Language Models

Lookup NU author(s): Professor Wei PanORCiD

Downloads

Full text for this publication is not currently held within this repository. Alternative links are provided below where available.


Abstract

© 2004-2012 IEEE.Interactive navigation requires robots to actively modify cluttered environments to create traversable paths, going beyond passive obstacle avoidance. However, existing methods either depend on global maps and lack the reasoning capabilities to make interaction decisions from local observations, or are restricted to interactions with simple geometric objects, limiting their applicability in partially observable, unstructured environments. To address these challenges, we propose counterfactual interactive navigation, named CoIN, a vision-language model (VLM)-based hierarchical framework that integrates high-level interaction reasoning with low-level loco-manipulation policies for diverse objects. Specifically, we propose CoIN-VLM, a VLM that internalizes counterfactual reasoning to evaluate the effect of object removal on goal reachability, thereby deciding when interaction is necessary and which object to interact with. To further align such reasoning with the robot's physical capabilities, we inject robot skill descriptions into the VLM context and ground them into a metric-scale environmental representation, ensuring that the generated plans remain physically feasible. To execute the generated high-level plans, we develop a comprehensive skill library through reinforcement learning, specifically introducing traversability-oriented strategies to manipulate diverse objects for path clearance. Furthermore, a systematic benchmark in Isaac Sim is proposed to evaluate both the reasoning and execution aspects of interactive navigation. Extensive simulations and real-world experiments demonstrate that CoIN significantly outperforms representative baselines, achieving a 17% higher overall success rate and over 80% improvement in complex long-horizon scenarios compared to the best-performing baseline, while exhibiting robust generalization across diverse object categories. Our project page is available at https://coins-internav.github.io/ Note to Practitioners - This work addresses the practical challenge of enabling autonomous robots to reach goals in cluttered indoor environments where the path is blocked by movable objects, without relying on a global map. The primary application is warehouse automation, facility inspection, and disaster-response robots that must decide when to interact, which object should be moved, and how to interact with diverse objects. The fine-tuned vision-language reasoning module determines the timing of interaction and selects the object whose removal is most likely to open a useful path. The learned skill library then executes efficient physical interactions, such as pushing obstacles or opening doors, to create traversable space for navigation. This improves navigation efficiency and reliability by reducing unnecessary detours and allowing the robot to complete tasks that are infeasible for passive obstacle avoidance.


Publication metadata

Author(s): Zhou K, Wen Z, Zhuo Z, Yan Z, Wu P, Ieng Hou U, Li S, Gao H, Ding K, Cao W, Pan W, Liu C

Publication type: Article

Publication status: Published

Journal: IEEE Transactions on Automation Science and Engineering

Year: 2026

Volume: 23

Pages: 14881-14899

Online publication date: 17/08/2026

Acceptance date: 06/08/2026

ISSN (print): 1545-5955

ISSN (electronic): 1558-3783

Publisher: Institute of Electrical and Electronics Engineers Inc.

URL: https://doi.org/10.1109/TASE.2026.3724341

DOI: 10.1109/TASE.2026.3724341


Altmetrics

Altmetrics provided by Altmetric


Share