A united visual representation is a shared feature encoding framework in multimodal artificial intelligence that maps disparate visual modalities, such as static images and temporal video sequences, into a single cohesive embedding space. Rather than processing images and videos through isolated representation pipelines or separate projection mechanisms, this approach harmonizes visual inputs into a unified tokenization structure before interfacing with downstream architectures such as large language models. By establishing an integrated feature representation across differing visual formats, it facilitates joint learning and cross-modal alignment, allowing computational models to process both static spatial details and dynamic temporal motions within a common semantic space.