// 01
Leveraging Large Vision-Language Models (LVLMs) for embodied reasoning and decision-making.
// 02
Instruction following and long-horizon planning in complex environments (VLN).
// 03
Parameter-efficient fine-tuning (PEFT) and alignment of foundation models for robotics tasks.
// 04
Bridging the gap between high-level linguistic instructions and low-level control policies.
// 05
Cross-modal representation learning for embodied agents in 3D environments.