[KDD'2026] "VideoRAG: Chat with Your Videos"
-
Updated
Mar 18, 2026 - Python
[KDD'2026] "VideoRAG: Chat with Your Videos"
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.
An open-weight 11B model series for long-form and real-time video understanding
[CVPR 2024] MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
Official code for Goldfish model for long video understanding and MiniGPT4-video for short video understanding
✨✨[NeurIPS 2025] This is the official implementation of our paper "Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension"
🔥🔥MLVU: Multi-task Long Video Understanding Benchmark
[CVPR 2026] LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Multi-granularity Correspondence Learning from Long-term Noisy Videos [ICLR 2024, Oral]
Context-Aware RAG library for Knowledge Graph ingestion and retrieval functions.
OmniAgent (ICML 2026): the first native omni-modal agent for active video perception — a 7B agent that beats Qwen2.5-VL-72B with 73% fewer frames on LVBench.
[ICLR 2025] TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos. (CVPR 2025))
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
[EMNLP 2023] TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding
This is the official implementation of ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory
Code for our ACL 2025 paper "Language Repository for Long Video Understanding"
[CVPR 2026] Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
[NeurIPS 2025] HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models
To associate your repository with the long-video-understanding topic, visit your repo's landing page and select "manage topics."