keyword
masked auto-encoder pre-training
Masked auto-encoder pre-training is a self-supervised learning paradigm in which a neural network learns generalized data representations by reconstructing missing or masked portions of an input from the visible parts. Under this framework, a subset of input elements, such as image patches or textual tokens, is intentionally masked out. An encoder transforms the visible or partially corrupted input into latent semantic embeddings, after which a decoder reconstructs the original, uncorrupted data using those embeddings. By requiring the model to infer hidden content from contextual cues, masked auto-encoder pre-training effectively captures high-level semantics, global dependencies, and underlying structures, producing robust foundational representations that can be fine-tuned for diverse downstream tasks across computer vision and natural language processing.
1 item

