// 01

Leveraging Large Vision-Language Models (LVLMs) for embodied reasoning and decision-making.

// 02

Instruction following and long-horizon planning in complex environments (VLN).

// 03

Parameter-efficient fine-tuning (PEFT) and alignment of foundation models for robotics tasks.

// 04

Bridging the gap between high-level linguistic instructions and low-level control policies.

// 05

Cross-modal representation learning for embodied agents in 3D environments.