Research

Paper

AI LLM March 12, 2026

Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing

Authors

Baifeng Shi, Stephanie Fu, Long Lian, Hanrong Ye, David Eigen, Aaron Reite, Boyi Li, Jan Kautz, Song Han, David M. Chan, Pavlo Molchanov, Trevor Darrell, Hongxu Yin

Abstract

Multi-modal large language models (MLLMs) have advanced general-purpose video understanding but struggle with long, high-resolution videos -- they process every pixel equally in their vision transformers (ViTs) or LLMs despite significant spatiotemporal redundancy. We introduce AutoGaze, a lightweight module that removes redundant patches before processed by a ViT or an MLLM. Trained with next-token prediction and reinforcement learning, AutoGaze autoregressively selects a minimal set of multi-scale patches that can reconstruct the video within a user-specified error threshold, eliminating redundancy while preserving information. Empirically, AutoGaze reduces visual tokens by 4x-100x and accelerates ViTs and MLLMs by up to 19x, enabling scaling MLLMs to 1K-frame 4K-resolution videos and achieving superior results on video benchmarks (e.g., 67.0% on VideoMME). Furthermore, we introduce HLVid: the first high-resolution, long-form video QA benchmark with 5-minute 4K-resolution videos, where an MLLM scaled with AutoGaze improves over the baseline by 10.1% and outperforms the previous best MLLM by 4.5%. Project page: https://autogaze.github.io/.

Metadata

arXiv ID: 2603.12254

Provider: ARXIV

Primary Category: cs.CV

Published: 2026-03-12

Fetched: 2026-03-14 05:03

Related papers

Gen-Searcher: Reinforcing Agentic Search for Image Generation

Kaituo Feng, Manyuan Zhang, Shuang Chen, Yunlong Lin, Kaixuan Fan, Yilei Jian... • 2026-03-30

On-the-fly Repulsion in the Contextual Space for Rich Diversity in Diffusion Transformers

Omer Dahary, Benaya Koren, Daniel Garibi, Daniel Cohen-Or • 2026-03-30

Graphilosophy: Graph-Based Digital Humanities Computing with The Four Books

Minh-Thu Do, Quynh-Chau Le-Tran, Duc-Duy Nguyen-Mai, Thien-Trang Nguyen, Khan... • 2026-03-30

ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining

Anuj Diwan, Eunsol Choi, David Harwath • 2026-03-30

RAD-AI: Rethinking Architecture Documentation for AI-Augmented Ecosystems

Oliver Aleksander Larsen, Mahyar T. Moghaddam • 2026-03-30

Raw Data (Debug)

{
  "raw_xml": "<entry>\n    <id>http://arxiv.org/abs/2603.12254v1</id>\n    <title>Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing</title>\n    <updated>2026-03-12T17:58:52Z</updated>\n    <link href='https://arxiv.org/abs/2603.12254v1' rel='alternate' type='text/html'/>\n    <link href='https://arxiv.org/pdf/2603.12254v1' rel='related' title='pdf' type='application/pdf'/>\n    <summary>Multi-modal large language models (MLLMs) have advanced general-purpose video understanding but struggle with long, high-resolution videos -- they process every pixel equally in their vision transformers (ViTs) or LLMs despite significant spatiotemporal redundancy. We introduce AutoGaze, a lightweight module that removes redundant patches before processed by a ViT or an MLLM. Trained with next-token prediction and reinforcement learning, AutoGaze autoregressively selects a minimal set of multi-scale patches that can reconstruct the video within a user-specified error threshold, eliminating redundancy while preserving information. Empirically, AutoGaze reduces visual tokens by 4x-100x and accelerates ViTs and MLLMs by up to 19x, enabling scaling MLLMs to 1K-frame 4K-resolution videos and achieving superior results on video benchmarks (e.g., 67.0% on VideoMME). Furthermore, we introduce HLVid: the first high-resolution, long-form video QA benchmark with 5-minute 4K-resolution videos, where an MLLM scaled with AutoGaze improves over the baseline by 10.1% and outperforms the previous best MLLM by 4.5%. Project page: https://autogaze.github.io/.</summary>\n    <category scheme='http://arxiv.org/schemas/atom' term='cs.CV'/>\n    <published>2026-03-12T17:58:52Z</published>\n    <arxiv:comment>CVPR 2026. Project page: https://autogaze.github.io/</arxiv:comment>\n    <arxiv:primary_category term='cs.CV'/>\n    <author>\n      <name>Baifeng Shi</name>\n    </author>\n    <author>\n      <name>Stephanie Fu</name>\n    </author>\n    <author>\n      <name>Long Lian</name>\n    </author>\n    <author>\n      <name>Hanrong Ye</name>\n    </author>\n    <author>\n      <name>David Eigen</name>\n    </author>\n    <author>\n      <name>Aaron Reite</name>\n    </author>\n    <author>\n      <name>Boyi Li</name>\n    </author>\n    <author>\n      <name>Jan Kautz</name>\n    </author>\n    <author>\n      <name>Song Han</name>\n    </author>\n    <author>\n      <name>David M. Chan</name>\n    </author>\n    <author>\n      <name>Pavlo Molchanov</name>\n    </author>\n    <author>\n      <name>Trevor Darrell</name>\n    </author>\n    <author>\n      <name>Hongxu Yin</name>\n    </author>\n  </entry>"
}