Structured-Noise Masked Modeling for Video, Audio and Beyond

Abstract

We introduce a structured-noise approach to masked modeling that generalizes across video, audio, and other modalities, improving self-supervised representation learning. This work was selected for oral presentation at ECCV 2026.

Publication
In 2026 European Conference on Computer Vision
Self-Supervised Learning Video Understanding Audio Masked Modeling Computer Vision