Open-category human-object interaction refers to a computer vision paradigm in which a system detects, localizes, and recognizes relationships between people and objects without being restricted to a fixed, predefined set of interaction classes. In standard human-object interaction tasks, models are limited to identifying specific human, action, and object triplets encountered during training. By contrast, open-category approaches enable the detection of novel, unseen, or open-vocabulary interactions in unconstrained environments by leveraging multimodal representations and natural language descriptions, such as phrases or interpretive sentences, facilitating zero-shot recognition and broader generalization in real-world visual understanding.