We introduce a structured-noise approach to masked modeling that generalizes across video, audio, and other modalities, improving self-supervised representation learning. This work was selected for oral presentation at ECCV 2026.
We explore data-independent masking strategies for Masked AutoEncoders (MAE), showing that carefully designed structured noise masks can match or improve upon learned/data-dependent masking, while being simpler and more efficient. See our [Project …