Project Details
Grant Program
Faculty Development Competitive Research Grants Program for 2026-2028
Project Description
This project aims to develop an efficient large language model (LLM)-based framework for aerial video understanding (AVU) to imporve intelligent surveillance. While unmanned aerial vehicles (UAVs) provide flexible and scalable monitoring, existing deep learning methods, mainly CNNs and Transformers, struggle with domain generalization, temporal modeling, and semantic reasoning. To address these challenges, the proposed research will adapt LLMs for multimodal aerial surveillance by aligning video and textual information to enable cross-modal reasoning and interpretable situational awareness. The project will bridge the domain gap through specialized dataset curation, design temporal adapters to capture motion dynamics, and develop lightweight architectures for efficient edge deployment. Furthermore, it will integrate aerial and ground-based video streams to achieve unified, real-time, and semantically rich event understanding. Overall, the proposed framework aims to deliver deployable and resource-efficient video–language models that advance the scalability and effectiveness of modern surveillance systems.
| Status | Active |
|---|---|
| Effective start/end date | 4/1/26 → 12/31/28 |
Keywords
- Video understanding
- Aerial surveillance
- Multimodal learning
- Language modeling
- Dataset curation
- Vision-language model
- Remote sensing
- Edge intelligence
Fingerprint
Explore the research topics touched on by this project. These labels are generated based on the underlying awards/grants. Together they form a unique fingerprint.