We filter white noise into different color noise distributions to build structured masks for masked video and audio modeling: Green3D for video and Regularized Blue Noise for audio, improving over random masking at no additional computational cost.
We explore data-independent masking strategies for Masked AutoEncoders (MAE), showing that carefully designed structured noise masks can match or improve upon learned/data-dependent masking, while being simpler and more efficient. See our [Project …