Video Understanding

Structured-Noise Masked Modeling for Video, Audio and Beyond

We introduce a structured-noise approach to masked modeling that generalizes across video, audio, and other modalities, improving self-supervised representation learning. This work was selected for oral presentation at ECCV 2026.