AgentVLN: Enhancing Vision Language Navigation with a reasoning Agent based on Explicit memory module
Undergraduate Thesis Completed · Publication in Preparation, 2026
Current vision-language models often lack structural scene understanding and fine-grained details information. These VLNs typically do not maintain an explicit structured memory module to store and organize previously observed information. This limitation becomes especially pronounced in long-horizon tasks. In this work, we propose a novel pipeline that constructs an open-vocabulary,real-time scene-graph from RGBD inputs, video captions and other multimodal inferred information. Our method further clusters scene-graph objects and regions in real-time. Building on this representation, we introduce an LLM-based planning agent that reasons over the continuously updated scene graph to decompose long-horizon instructions, and automatically detect deviation from goal and perform corrective replanning. The agent then publishes more detailed, short-term sub-instructions to a downstream VLN model instead of original long-horizon instructions. We also developed RL-based Husky Locomotion module for downstream of VLN and plan to test in real world scenarios
During student researcher at Toronto Robotics and AI Lab