Incorporating Hierarchy into Text Encoder: a Contrastive Learning Approach for Hierarchical Text Classification
Zihan WangPeiyi WangLianzhe HuangXin SunHoufeng Wang
Proposes a hierarchy-guided contrastive learning framework that directly embeds taxonomic label relationships into the text encoder using modified Graphormer structures, removing the need for separate, redundant label representations during inference.
Hierarchical text classification categorizes documents into structured, tree-like label taxonomies, which is critical for organizing complex digital content in news publishing, academic archiving, and enterprise knowledge management. Conventional approaches typically rely on two separate neural networks—one for the text and one for the label hierarchy—and fuse their outputs during classification. However, because the structural hierarchy remains identical across all inputs, passing this static graph through a secondary model during inference adds computational redundancy without providing dynamic, context-specific representations.
The article demonstrates that taxonomy structures can be directly embedded into a standard text encoder during training through contrastive learning, completely eliminating the need for a separate graph network during deployment. The authors evaluate this framework, named Hierarchy-Guided Contrastive Learning, across multiple standard benchmark datasets against prevailing state-of-the-art architectures.
The approach uses a graph transformer to encode structural relationships and natural language label names during training. This structure guides the construction of positive training examples by identifying and retaining key task-relevant words while masking non-essential text. The main text encoder is then trained to pull matching pairs closer together in representation space while pushing unrelated samples apart. The authors tested this method across three large public benchmark corpora spanning academic papers and news articles: Web of Science (over 46,000 samples), NYTimes (over 36,000 samples), and RCV1-V2 (over 800,000 samples).
The evaluation produced several notable results. On the Web of Science benchmark, the framework improved performance over the baseline text model by 1.5% in Micro-F1 (reaching 87.11%) and 2.1% in Macro-F1 (reaching 81.20%), establishing new performance highs over existing baselines. On the NYTimes corpus, it achieved a 2.3% boost in Macro-F1 over standard text modeling. Component analyses revealed that semantic label names were the single most important element in the structural graph encoder; removing them caused the steepest drop in classification accuracy. Furthermore, hierarchy-guided sample generation outperformed alternative techniques, such as random word masking and gradient-based adversarial perturbations.
These findings indicate that organizations can improve classification accuracy on complex taxonomy structures without incurring additional computational latency during production deployment. Because the structural hierarchy is permanently embedded into the primary text encoder, the graph-processing component can be discarded after training. This reduces inference overhead and infrastructure complexity for production systems.
Technical leaders and practitioners working with hierarchical categorization should consider incorporating hierarchy-guided contrastive objectives into their pre-deployment training pipelines. When applying this technique, organizations must ensure that label taxonomies feature meaningful, descriptive text names rather than arbitrary numeric or coded identifiers, as semantic label embeddings are vital to achieving performance gains. Further pilot validation is recommended when evaluating datasets lacking descriptive category names.
- Paper: SimCSE: Simple Contrastive Learning of Sentence Embeddings, Tianyu Gao et al. (2021). SimCSE introduces the core contrastive sentence embedding framework that the source paper builds upon and modifies with hierarchy-guided data augmentations.
- Paper: Graph Contrastive Learning with Adaptive Augmentation, Yanqiao Zhu et al. (2020). This paper establishes adaptive augmentation principles for contrastive learning on graphs, providing essential background for how structural importance can guide sample corruption.
- Paper: Poincaré Embeddings for Learning Hierarchical Representations, Maximilian Nickel et al. (2017). Reading this paper introduces foundational principles of embedding hierarchical tree structures and taxonomies into vector representations.
- Paper: Support vector machine learning for interdependent and structured output spaces, Ioannis Tsochantaridis et al. (2004). This foundational work formalizes classification over structured, interdependent output spaces and hierarchical taxonomies.
- Paper: Graph Transformer Networks, Seongjun Yun et al. (2019). This work introduces graph transformer networks, foundational to how the source paper's training architecture encodes label hierarchies and structural relationships.
- Paper: Graph Convolutional Networks for Text Classification, Liang Yao et al. (2018). This paper establishes the standard baseline paradigm of using graph neural networks for text classification that the source paper seeks to simplify during inference.
- Paper: HiCLIP: Contrastive Language-Image Pretraining with Hierarchy-aware Attention, Shijie Geng et al. (2023). HiCLIP extends hierarchy-aware contrastive representation learning from text classification to multimodal vision-language pretraining architectures.
- Paper: RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank, Jiduan Liu et al. (2023). RankCSE generalizes contrastive sentence representation learning by incorporating fine-grained ranking objectives rather than binary positive-negative pairings.
- Paper: MA-GCL: Model Augmentation Tricks for Graph Contrastive Learning, Xumeng Gong et al. (2023). MA-GCL builds on graph contrastive learning by exploring architectural view augmentations rather than input masking perturbations.
- Paper: Attribute and Structure Preserving Graph Contrastive Learning, Jialu Chen et al. (2023). This paper extends graph contrastive learning techniques by developing multi-view objectives that preserve both local node attributes and higher-order structural topologies.
