The Mutual-Enhanced Incongruity Learning Network is a multimodal deep learning architecture designed for sarcasm detection that identifies sarcastic content by capturing discrepancies within and between different data modalities, such as text and images. The network operates by integrating a global incongruity learning module with a local semantic-guided incongruity learning module, allowing it to focus on fine-grained critical features while maintaining broader contextual awareness. By utilizing a mutual enhancement mechanism, the framework exploits the consistency between local and global representations to filter out irrelevant information, mitigate noise, and accurately model subtle semantic contradictions that characterize sarcasm.