Prompt2Fly
Using LLMs and VLMs for Autonomous Aerial Transportation
Aerial robots carrying cable-suspended payloads provide a lightweight and mechanically simple solution for time-critical delivery in unstructured and difficult-to-access environments. However, their underactuated and coupled dynamics make autonomous payload pickup and transportation using only onboard sensing particularly challenging. To address this gap, we present Prompt2Fly, a hierarchical framework that combines large language models (LLMs) and vision-language models (VLMs) to translate natural-language commands into executable aerial payload-transportation missions. Prompt2Fly uses an LLM to generate reactive behavior trees whose nodes are grounded in validated robot capabilities. Each behavior tree decomposes the commanded task into a sequence of interpretable behaviors and supplies task-specific language context to downstream VLM-based perception and planning modules. This architecture combines the semantic reasoning and generalization capabilities of pretrained models with the reliability, interpretability, and modularity of established planning and control algorithms. We validate Prompt2Fly through extensive simulation and real-world experiments using multiple onboard cameras and a magnetic pickup mechanism. The system autonomously identifies, picks up, transports, and delivers payloads across diverse environments and task specifications. To the best of our knowledge, Prompt2Fly is the first system to demonstrate onboard vision-based autonomous pickup and transportation using a cable-suspended aerial robot in both simulation and real-world experiments.