Submitted successfully

Track Ⅶ

Deep Models for Aerial Robotics: Perception, Reasoning, and Edge Computing    (Submission Deadline: November 15, 2026)
面向空中机器人的感知推理与边缘计算深度模型
     
Chair:  Co-chair:   
Lingyan Ran Xiaoqiang Zhang
Northwestern Polytechnical University, China Southwest University of Science and Technology, China
Topics:  
  • The adaptation of deep models for aerial imagery (深度模型(Deep Models)在航拍图像中的适配技术)
  • The design of lightweight networks for onboard processing (面向机载处理的轻量化网络设计)
  • The development of embodied AI systems that enable UAVs to perceive, reason, and act autonomously. (赋能无人机实现自主感知、推理与行动的具身智能(Embodied AI)系统开发)
   
Summary:  

Unmanned Aerial Vehicles (UAVs) are rapidly evolving from remotely controlled sensors into autonomous intelligent agents. While traditional visual perception algorithms (detection, tracking, and segmentation) have reached a level of maturity, the next frontier lies in semantic understanding and real-time reasoning in complex, open-world environments.
The emergence of Vision-Language Models (VLMs) and Large Multimodal Models (LMMs) offers unprecedented capabilities for zero-shot learning, open-vocabulary detection, and human-machine interaction. However, deploying these massive models on UAVs presents significant challenges due to the strict constraints on Size, Weight, and Power (SWaP).
This Session aims to bridge the gap between high-level artificial intelligence (AI) capabilities and low-level edge implementation in aerial robotics. 

   
无人机(UAV)正迅速从遥控感应设备演变为自主智能体。尽管传统的视觉感知算法(检测、跟踪和分割)已日趋成熟,但下一个技术前沿在于复杂开放世界环境中的语义理解与实时推理。
视觉语言模型(VLM)和多模态大模型(LMM)的出现,为零样本学习、开放词汇检测以及人机交互带来了前所未有的能力。然而,受限于无人机在尺寸、重量与功耗(SWaP)方面的严格约束,在机载设备上部署这些庞大的模型仍面临着重大挑战。
本专题旨在填补空中机器人领域中高层人工智能(AI)能力与底层边缘端实现之间的鸿沟。