Skip to main navigation Skip to search Skip to main content

Advancing Aerial Surveillance through Efficient Video–Language Models

Project: FDCRGP

Project Details

Grant Program

Faculty Development Competitive Research Grants Program for 2026-2028

Project Description

This project aims to develop an efficient large language model (LLM)-based framework for aerial video understanding (AVU) to imporve intelligent surveillance. While unmanned aerial vehicles (UAVs) provide flexible and scalable monitoring, existing deep learning methods, mainly CNNs and Transformers, struggle with domain generalization, temporal modeling, and semantic reasoning. To address these challenges, the proposed research will adapt LLMs for multimodal aerial surveillance by aligning video and textual information to enable cross-modal reasoning and interpretable situational awareness. The project will bridge the domain gap through specialized dataset curation, design temporal adapters to capture motion dynamics, and develop lightweight architectures for efficient edge deployment. Furthermore, it will integrate aerial and ground-based video streams to achieve unified, real-time, and semantically rich event understanding. Overall, the proposed framework aims to deliver deployable and resource-efficient video–language models that advance the scalability and effectiveness of modern surveillance systems.
StatusActive
Effective start/end date4/1/2612/31/28

Keywords

  • Video understanding
  • Aerial surveillance
  • Multimodal learning
  • Language modeling
  • Dataset curation
  • Vision-language model
  • Remote sensing
  • Edge intelligence

Fingerprint

Explore the research topics touched on by this project. These labels are generated based on the underlying awards/grants. Together they form a unique fingerprint.