OMG-LLaVA is a multimodal artificial intelligence framework designed to bridge image-level conversational reasoning, object-level identification, and pixel-level visual understanding within a single system. Unlike traditional multimodal large language models that primarily handle high-level visual descriptions or specialized computer vision tools limited to segmentation tasks, OMG-LLaVA integrates a universal segmentation encoder and decoder with a large language model. This unified architecture processes user text instructions alongside diverse visual prompts by converting perception priors and visual features into tokens that the language model can interpret. Consequently, the system can engage in complex visual reasoning and concurrently produce conversational text responses and precise pixel-level segmentation masks.