Incorporating Hierarchy into Text Encoder: a Contrastive Learning Approach for Hierarchical Text Classification

Zihan WangPeiyi WangLianzhe HuangXin SunHoufeng Wang

article2022ACL153 citations

Proposes a hierarchy-guided contrastive learning framework that directly embeds taxonomic label relationships into the text encoder using modified Graphormer structures, removing the need for separate, redundant label representations during inference.

Listen

Hierarchical text classification categorizes documents into structured, tree-like label taxonomies, which is critical for organizing complex digital content in news publishing, academic archiving, and enterprise knowledge management. Conventional approaches typically rely on two separate neural networks—one for the text and one for the label hierarchy—and fuse their outputs during classification. However, because the structural hierarchy remains identical across all inputs, passing this static graph through a secondary model during inference adds computational redundancy without providing dynamic, context-specific representations.

The article demonstrates that taxonomy structures can be directly embedded into a standard text encoder during training through contrastive learning, completely eliminating the need for a separate graph network during deployment. The authors evaluate this framework, named Hierarchy-Guided Contrastive Learning, across multiple standard benchmark datasets against prevailing state-of-the-art architectures.

The approach uses a graph transformer to encode structural relationships and natural language label names during training. This structure guides the construction of positive training examples by identifying and retaining key task-relevant words while masking non-essential text. The main text encoder is then trained to pull matching pairs closer together in representation space while pushing unrelated samples apart. The authors tested this method across three large public benchmark corpora spanning academic papers and news articles: Web of Science (over 46,000 samples), NYTimes (over 36,000 samples), and RCV1-V2 (over 800,000 samples).

The evaluation produced several notable results. On the Web of Science benchmark, the framework improved performance over the baseline text model by 1.5% in Micro-F1 (reaching 87.11%) and 2.1% in Macro-F1 (reaching 81.20%), establishing new performance highs over existing baselines. On the NYTimes corpus, it achieved a 2.3% boost in Macro-F1 over standard text modeling. Component analyses revealed that semantic label names were the single most important element in the structural graph encoder; removing them caused the steepest drop in classification accuracy. Furthermore, hierarchy-guided sample generation outperformed alternative techniques, such as random word masking and gradient-based adversarial perturbations.

These findings indicate that organizations can improve classification accuracy on complex taxonomy structures without incurring additional computational latency during production deployment. Because the structural hierarchy is permanently embedded into the primary text encoder, the graph-processing component can be discarded after training. This reduces inference overhead and infrastructure complexity for production systems.

Technical leaders and practitioners working with hierarchical categorization should consider incorporating hierarchy-guided contrastive objectives into their pre-deployment training pipelines. When applying this technique, organizations must ensure that label taxonomies feature meaningful, descriptive text names rather than arbitrary numeric or coded identifiers, as semantic label embeddings are vital to achieving performance gains. Further pilot validation is recommended when evaluating datasets lacking descriptive category names.

arXiv: 2203.03825
Cover for Incorporating Hierarchy into Text Encoder: a Contrastive Learning Approach for Hierarchical Text Classification

Abstract

Hierarchical text classification is a challenging subtask of multi-label classification due to its complex label hierarchy. Existing methods encode text and label hierarchy separately and mix their representations for classification, where the hierarchy remains unchanged for all input text. Instead of modeling them separately, in this work, we propose Hierarchy-guided Contrastive Learning (HGCLR) to directly embed the hierarchy into a text encoder. During training, HGCLR constructs positive samples for input text under the guidance of the label hierarchy. By pulling together the input text and its positive sample, the text encoder can learn to generate the hierarchy-aware text representation independently. Therefore, after training, the HGCLR enhanced text encoder can dispense with the redundant hierarchy. Extensive experiments on three benchmark datasets verify the effectiveness of HGCLR.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Hierarchical Text Classification
  • 2.2 Contrastive Learning
  • 3 Problem Definition
  • 4 Methodology
  • 4.1 Text Encoder
  • 4.2 Graph Encoder
  • 4.3 Positive Sample Generation
  • 4.4 Contrastive Learning Module
  • 4.5 Classification and Objective Function
  • 5 Experiments
  • 5.1 Experiment Setup
  • 5.2 Experimental Results
  • 5.3 Analysis
  • 5.3.1 Effect of Hierarchy
  • 5.3.2 Effect of Graphormer
  • 5.3.3 Effect of Positive Example Generation
  • 6 Conclusion
  • Acknowledgements
  • References
  • A Trick for Token Selection

Knowls

  1. Knowl 1 — Hierarchy-Guided Contrastive Learning for HTC

    model/method

    Hierarchy-Guided Contrastive Learning (HGCLR) injects a taxonomic label hierarchy into a BERT text encoder during training rather than combining separate text and hierarchy representations at classification time. A customized Graphormer encodes the label hierarchy and produces label features. For each labeled input text, HGCLR uses the label features to retain tokens that are likely to express the ground-truth labels, creating a shortened positive sequence. The original sequence and its shortened counterpart are processed by the same BERT encoder; their representations are pulled together with contrastive learning, while representations from different examples in the minibatch are pushed apart. Classification losses are applied to both sequences. At test time, HGCLR discards the graph encoder and positive-sample generator and classifies using only the hierarchy-aware BERT encoder and its classification head.

  2. Knowl 2 — Hierarchical Text Classification Problem Formulation

    definition

    For an input token sequence x={x1,…,xn}x=\{x_1,\ldots,x_n\}, hierarchical text classification predicts a subset y⊆Yy\subseteq Y of kk predefined labels. The labels are nodes of a directed acyclic graph G=(Y,E)G=(Y,E), where EE represents taxonomic relations. In the datasets used by HGCLR, every non-root label has exactly one parent, so the hierarchy can be treated as tree-like. The valid label set for a document is path-consistent: if a non-root label yj∈yy_j\in y is assigned, its parent label must also belong to yy. WOS uses a single hierarchical path per example, whereas NYT and RCV1-V2 can contain labels from multiple paths.

  3. Knowl 3 — Customized Graphormer for Label-Hierarchy Encoding

    model/method

    HGCLR represents every label node yiy_i with a vector fi∈Rdhf_i\in\mathbb{R}^{d_h} formed by adding a learnable label embedding to the average of the BERT token embeddings of the label name. The label-name embedding shares weights with the text encoder. Stacking the kk node vectors gives F∈Rk×dhF\in\mathbb{R}^{k\times d_h}. A Graphormer-style self-attention layer modifies attention using both graph distance and edge information:

    AijG=(fiWQG)(fjWKG)⊤dh+cij+bϕ(yi,yj).A^G_{ij}=\frac{(f_iW_Q^G)(f_jW_K^G)^\top}{\sqrt{d_h}}+c_{ij}+b_{\phi(y_i,y_j)}.

    Here WQG,WKG∈Rdh×dhW_Q^G,W_K^G\in\mathbb{R}^{d_h\times d_h} are learnable projections, ϕ(yi,yj)\phi(y_i,y_j) is the shortest-path distance between labels yiy_i and yjy_j, and bϕ(yi,yj)b_{\phi(y_i,y_j)} is a learnable scalar for that distance. If the path between the two labels contains edges e1,…,eDe_1,\ldots,e_D, the edge term is cij=D−1∑r=1Dwerc_{ij}=D^{-1}\sum_{r=1}^{D}w_{e_r}, where each werw_{e_r} is a learnable scalar. The graph-aware attention output is

    L=LayerNorm⁡ ⁣(softmax⁡(AG)V+F),L=\operatorname{LayerNorm}\!\left(\operatorname{softmax}(A^G)V+F\right),

    where VV is the value matrix in the attention layer and L∈Rk×dhL\in\mathbb{R}^{k\times d_h} contains the label features used for token selection. Unlike a local graph-attention layer, this Graphormer permits every label node to attend globally to all other labels.

  4. Knowl 4 — Hierarchy-Guided Positive-Sample Generation

    algorithm

    For each labeled input sequence, HGCLR constructs a shortened positive sample by estimating how strongly each token expresses each ground-truth label. Let ei∈Rdhe_i\in\mathbb{R}^{d_h} be the BERT embedding of token xix_i, let lj∈Rdhl_j\in\mathbb{R}^{d_h} be the graph-encoder feature for label yjy_j, and let WQ,WK∈Rdh×dhW_Q,W_K\in\mathbb{R}^{d_h\times d_h} be learnable projections. The token-label attention score is

    qi=eiWQ,kj=ljWK,Aij=qikj⊤dh.q_i=e_iW_Q,\qquad k_j=l_jW_K,\qquad A_{ij}=\frac{q_i k_j^\top}{\sqrt{d_h}}.

    For each token, a Gumbel-Softmax over the kk labels gives differentiable label-assignment probabilities Pij=gumbel_softmax⁡(Ai1,…,Aik)jP_{ij}=\operatorname{gumbel\_softmax}(A_{i1},\ldots,A_{ik})_j. If yy is the ground-truth label set, the token's total ground-truth relevance is Pi=∑j∈yPijP_i=\sum_{j\in y}P_{ij}. Given a threshold γ\gamma, the positive sequence keeps token xix_i when Pi>γP_i>\gamma and replaces it otherwise with a zero embedding that preserves the token position:

    x^i={xi,Pi>γ,0,Pi≤γ.\hat{x}_i=\begin{cases}x_i,&P_i>\gamma,\\0,&P_i\leq\gamma.\end{cases}

    The same BERT encoder processes the original sequence and the resulting shortened sequence, producing sentence representations hxh_x and h^x\hat h_x from their first-token states. Because the retained tokens are selected using both the label hierarchy and the ground-truth labels, the positive sample is intended to preserve the document's classification labels while removing less informative content.

  5. Knowl 5 — Differentiable Implementation of Token Selection

    model/method

    The hard threshold in HGCLR's token selector is implemented on token embeddings so that the selector can still receive gradients. For token embedding eie_i and relevance probability PiP_i, the positive embedding is

    e^i=ei({Pi+Detach⁡(1−Pi),Pi>γ,0,Pi≤γ,)\hat e_i=e_i\left(\begin{cases}P_i+\operatorname{Detach}(1-P_i),&P_i>\gamma,\\0,&P_i\leq\gamma,\end{cases}\right)

    where Detach⁡(u)\operatorname{Detach}(u) has numerical value uu but blocks gradients through uu. Thus the retained embedding is numerically equal to eie_i, while the discarded embedding is numerically zero. The derivative with respect to the relevance probability is

    ∂e^i∂Pi={ei,Pi>γ,0,Pi≤γ.\frac{\partial\hat e_i}{\partial P_i}=\begin{cases}e_i,&P_i>\gamma,\\0,&P_i\leq\gamma.\end{cases}

    This preserves the discrete keep-or-remove behavior in the forward pass while allowing back-propagation to update the token-label attention and the graph encoder.

  6. Knowl 6 — Contrastive Representation Learning Objective

    equation

    For a minibatch of NN original-positive pairs, let hi,h^i∈Rdhh_i,\hat h_i\in\mathbb{R}^{d_h} be the BERT sentence representations of the original and hierarchy-guided positive sequences. HGCLR applies the same two-layer projection head to both views:

    ci=W2ReLU⁡(W1hi),c^i=W2ReLU⁡(W1h^i),c_i=W_2\operatorname{ReLU}(W_1h_i),\qquad \hat c_i=W_2\operatorname{ReLU}(W_1\hat h_i),

    where W1,W2∈Rdh×dhW_1,W_2\in\mathbb{R}^{d_h\times d_h}. Let Z={c1,…,cN,c^1,…,c^N}Z=\{c_1,\ldots,c_N,\hat c_1,\ldots,\hat c_N\} contain the 2N2N projected representations. The positive partner of cic_i is c^i\hat c_i, and the positive partner of c^i\hat c_i is cic_i; all other representations in the minibatch are negatives. With cosine similarity sim⁡(u,v)=u⊤v/(∥u∥∥v∥)\operatorname{sim}(u,v)=u^\top v/(\lVert u\rVert\lVert v\rVert), temperature τ\tau, and matching function μ\mu, the NT-Xent loss for zm∈Zz_m\in Z is

    Lmcon=−log⁡exp⁡(sim⁡(zm,μ(zm))/τ)∑i=1, i≠m2Nexp⁡(sim⁡(zm,zi)/τ).\mathcal{L}^{\mathrm{con}}_m=-\log\frac{\exp(\operatorname{sim}(z_m,\mu(z_m))/\tau)}{\sum_{i=1,\,i\neq m}^{2N}\exp(\operatorname{sim}(z_m,z_i)/\tau)}.

    The batch contrastive loss is the mean over the 2N2N views:

    Lcon=12N∑m=12NLmcon.\mathcal{L}^{\mathrm{con}}=\frac{1}{2N}\sum_{m=1}^{2N}\mathcal{L}^{\mathrm{con}}_m.

    The objective makes the BERT representations of an input and its label-preserving shortened version similar, while separating representations belonging to different documents.

  7. Knowl 7 — Joint Classification and Training Objective

    equation

    HGCLR uses flat multi-label classification on top of the BERT sentence representation. For document ii, let hi∈Rdhh_i\in\mathbb{R}^{d_h} be the original-sequence representation, yij∈{0,1}y_{ij}\in\{0,1\} indicate whether label jj is assigned, and k=∣Y∣k=|Y| be the number of labels. The predicted probability is

    pij=sigmoid⁡(WChi+bC)j,p_{ij}=\operatorname{sigmoid}(W_Ch_i+b_C)_j,

    where WC∈Rk×dhW_C\in\mathbb{R}^{k\times d_h} and bC∈Rkb_C\in\mathbb{R}^{k}. The binary cross-entropy over a minibatch of NN documents is

    LC=∑i=1N∑j=1k[−yijlog⁡pij−(1−yij)log⁡(1−pij)].\mathcal{L}^{C}=\sum_{i=1}^{N}\sum_{j=1}^{k}\left[-y_{ij}\log p_{ij}-(1-y_{ij})\log(1-p_{ij})\right].

    An analogous loss L^C\hat{\mathcal{L}}^{C} is computed for the hierarchy-guided positive representations h^i\hat h_i. The complete HGCLR training loss is

    L=LC+L^C+λLcon,\mathcal{L}=\mathcal{L}^{C}+\hat{\mathcal{L}}^{C}+\lambda\mathcal{L}^{\mathrm{con}},

    where λ\lambda weights the contrastive term. At inference, only the original BERT representation and the sigmoid classification head are used; the graph encoder and positive-sample branch are not required.

  8. Knowl 8 — Datasets and Training Configuration

    experimental setup

    HGCLR was evaluated on Web-of-Science (WOS), NYTimes (NYT), and RCV1-V2 using Micro-F1 and Macro-F1. WOS contains scientific-paper abstracts and has single-path labels; NYT and RCV1-V2 are news datasets with multi-path taxonomies. Their statistics are:

    Could not parse LaTeX table

    The text encoder was bert-base-uncased; the customized Graphormer used hidden size dh=768d_h=768 and 8 attention heads. Training used batch size 12 and Adam with learning rate 3×10−53\times10^{-5}. Training stopped when development Macro-F1 failed to improve for 6 epochs. The token-selection threshold was γ=0.02\gamma=0.02 for WOS and γ=0.005\gamma=0.005 for both NYT and RCV1-V2. The contrastive-loss weight was λ=0.1\lambda=0.1 for WOS and RCV1-V2 and λ=0.3\lambda=0.3 for NYT; the contrastive temperature was fixed at τ=1\tau=1. The threshold and loss weight were selected by development-set grid search.

  9. Knowl 9 — Benchmark Performance of HGCLR

    data/table

    HGCLR was compared with hierarchy-aware classifiers and BERT-based variants under Micro-F1 and Macro-F1. The comparison tests whether injecting hierarchy into BERT during training improves over using BERT alone or combining BERT with a separate hierarchy module. HGCLR obtains the strongest WOS scores, the strongest NYT scores among the reported systems, and the strongest RCV1-V2 Micro-F1, although HiMatch has a higher RCV1-V2 Macro-F1. The reported results are:

    Could not parse LaTeX table

    Relative to the authors' BERT implementation, HGCLR improves WOS Micro-F1 from 85.63 to 87.11 and Macro-F1 from 79.07 to 81.20; NYT Macro-F1 from 65.62 to 67.96; and RCV1-V2 Micro-F1 from 85.65 to 86.49. RCV1-V2 provides no label names, so HGCLR cannot use its label-name embeddings there; the authors identify this as a likely reason its relative advantage is smaller on that dataset.

  10. Knowl 10 — Ablation Evidence for Graph Encoding and Contrastive Learning

    empirical result

    Component ablations on the WOS development set show that both graph encoding and contrastive learning contribute to HGCLR. The full system reaches Micro-F1 87.46 and Macro-F1 81.52, compared with 85.75 and 79.36 for BERT. Replacing the customized Graphormer with GCN or GAT reduces performance, and removing the graph encoder or contrastive loss produces larger drops:

    Could not parse LaTeX table

    Removing individual Graphormer components gives the following WOS development results:

    Could not parse LaTeX table

    These results indicate that label-name embeddings contribute most among the tested Graphormer components, followed by spatial encoding; edge encoding contributes least. The authors attribute the weak edge-encoding effect to the absence of informative edge features in the HTC taxonomies. A label-representation visualization also shows that labels sharing a parent become more clustered under HGCLR than under BERT, consistent with the intended hierarchy-aware representation.

  11. Knowl 11 — Hierarchy-Guided Sampling Beats Generic Positive-Sample Generation

    empirical result

    The positive-sample generator was compared with dropout, random token masking, and adversarial attack on the WOS development set. The methods were evaluated with approximately the same number of valid tokens retained for the random approaches. Hierarchy-guided sampling gives the best result on both metrics:

    Could not parse LaTeX table

    Removing only the contrastive loss still leaves the hierarchy-guided positive examples as a useful augmentation, achieving 86.72 Micro-F1 and 80.97 Macro-F1 rather than the BERT baseline's 85.75 and 79.36. Random masking performs better than dropout, suggesting that removing whole tokens creates harder contrastive examples than perturbing neuron activations. Adversarial examples perform worse because gradient-based perturbations are not constrained by the label hierarchy and therefore are not guaranteed to preserve the document's labels. The paper's qualitative examples show that tokens strongly associated with a ground-truth category, such as domain keywords, tend to be retained, while irrelevant function words and less label-informative content are removed.

Coverage note — No substantial contributed material was omitted; background, related work, acknowledgements, and illustrative reference material were excluded.

References

  1. 1.Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang. 2018. Generating natural language adversarial examples. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2890–2896, Brussels, Belgium. Association for Computational Linguistics.
  2. 2.Siddhartha Banerjee, Cem Akkaya, Francisco Perez-Sorrosal, and Kostas Tsioutsiouliklis. 2019. Hierarchical transfer learning for multi-label text classification. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6295–6300, Florence, Italy. Association for Computational Linguistics.
  3. 3.Boli Chen, Xin Huang, Lin Xiao, Zixin Cai, and Liping Jing. 2020a. Hyperbolic interaction model for hierarchical multi-label classification. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7496–7503.
  4. 4.Haibin Chen, Qianli Ma, Zhenxi Lin, and Jiangyue Yan. 2021. Hierarchy-aware label semantics matching network for hierarchical text classification. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4370–4379, Online. Association for Computational Linguistics.
  5. 5.Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020b. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR.
  6. 6.Zhongfen Deng, Hao Peng, Dongxiao He, Jianxin Li, and Philip Yu. 2021. HTCInfoMax: A global model for hierarchical text classification via information maximization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3259–3265, Online. Association for Computational Linguistics.
  7. 7.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  8. 8.Hongchao Fang, Sicheng Wang, Meng Zhou, Jiayuan Ding, and Pengtao Xie. 2020. Cert: Contrastive self-supervised learning for language understanding. arXiv preprint arXiv:2005.12766.
  9. 9.Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. Simcse: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6894–6910.
  10. 10.Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014. Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572.
  11. 11.Siddharth Gopal and Yiming Yang. 2013. Recursive regularization for large-scale classification with hierarchical and graphical dependencies. In Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 257–265.
  12. 12.Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9729–9738.
  13. 13.Eric Jang, Shixiang Gu, and Ben Poole. 2016. Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144.
  14. 14.Rie Johnson and Tong Zhang. 2015. Effective use of word order for text categorization with convolutional neural networks. In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 103–112, Denver, Colorado. Association for Computational Linguistics.
  15. 15.Taeuk Kim, Kang Min Yoo, and Sang-goo Lee. 2021. Self-guided contrastive learning for BERT sentence representations. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2528–2540, Online. Association for Computational Linguistics.
  16. 16.Kamran Kowsari, Donald E Brown, Mojtaba Heidarysafa, Kiana Jafari Meimandi, Matthew S Gerber, and Laura E Barnes. 2017. Hdltex: Hierarchical deep learning for text classification. In 2017 16th IEEE international conference on machine learning and applications (ICMLA), pages 364–371. IEEE.
  17. 17.Siwei Lai, Liheng Xu, Kang Liu, and Jun Zhao. 2015. Recurrent convolutional neural networks for text classification. In Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, pages 2267–2273.
  18. 18.David D Lewis, Yiming Yang, Tony Russell-Rose, and Fan Li. 2004. Rcv1: A new benchmark collection for text categorization research. Journal of machine learning research, 5(Apr):361–397.
  19. 19.Yuning Mao, Jingjing Tian, Jiawei Han, and Xiang Ren. 2019. Hierarchical text classification with reinforced label assignment. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 445–455, Hong Kong, China. Association for Computational Linguistics.
  20. 20.Yu Meng, Chenyan Xiong, Payal Bajaj, Paul Bennett, Jiawei Han, Xia Song, et al. 2021. Coco-lm: Correcting and contrasting text sequences for language model pretraining. Advances in Neural Information Processing Systems, 34.
  21. 21.Lin Pan, Chung-Wei Hang, Avirup Sil, Saloni Potdar, and Mo Yu. 2021. Improved text classification via contrastive adversarial training. arXiv preprint arXiv:2107.10137.
  22. 22.Hao Peng, Jianxin Li, Senzhang Wang, Lihong Wang, Qiran Gong, Renyu Yang, Bo Li, Philip Yu, and Lifang He. 2019. Hierarchical taxonomy-aware and attentional graph capsule rcnns for large-scale multi-label text classification. IEEE Transactions on Knowledge and Data Engineering.
  23. 23.Evan Sandhaus. 2008. The new york times annotated corpus. Linguistic Data Consortium, Philadelphia, 6(12):e26752.
  24. 24.Kazuya Shimura, Jiyi Li, and Fumiyo Fukumoto. 2018. HFT-CNN: Learning hierarchical category structure for multi-label short text categorization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 811–816, Brussels, Belgium. Association for Computational Linguistics.
  25. 25.Carlos N Silla and Alex A Freitas. 2011. A survey of hierarchical classification across different application domains. Data Mining and Knowledge Discovery, 22(1):31–72.
  26. 26.Samson Tan, Shafiq Joty, Min-Yen Kan, and Richard Socher. 2020. It’s morphin’ time! Combating linguistic discrimination with inflectional perturbations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2920–2935, Online. Association for Computational Linguistics.
  27. 27.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems, 30.
  28. 28.Boyan Wang, Xuegang Hu, Peipei Li, and S Yu Philip. 2021a. Cognitive structure learning model for hierarchical multi-label text classification. Knowledge-Based Systems, 218:106876.
  29. 29.Dong Wang, Ning Ding, Piji Li, and Haitao Zheng. 2021b. CLINE: Contrastive learning with semantic negative examples for natural language understanding. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 2332–2342, Online. Association for Computational Linguistics.
  30. 30.Jonatas Wehrmann, Ricardo Cerri, and Rodrigo Barros. 2018. Hierarchical multi-label classification networks. In International Conference on Machine Learning, pages 5075–5084. PMLR.
  31. 31.Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sholeifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online. Association for Computational Linguistics.
  32. 32.Jiawei Wu, Wenhan Xiong, and William Yang Wang. 2019. Learning to learn and predict: A meta-learning approach for multi-label classification. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4354–4364, Hong Kong, China. Association for Computational Linguistics.
  33. 33.Zhuofeng Wu, Sinong Wang, Jiatao Gu, Madian Khabsa, Fei Sun, and Hao Ma. 2020. Clear: Contrastive learning for sentence representation. arXiv preprint arXiv:2012.15466.
  34. 34.Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu. 2021. Do transformers really perform badly for graph representation? Advances in Neural Information Processing Systems, 34.
  35. 35.Wei Emma Zhang, Quan Z Sheng, Ahoud Alhazmi, and Chenliang Li. 2020. Adversarial attacks on deep-learning models in natural language processing: A survey. ACM Transactions on Intelligent Systems and Technology (TIST), 11(3):1–41.
  36. 36.Xinyi Zhang, Jiahao Xu, Charlie Soh, and Lihui Chen. 2021. La-hcn: Label-based attention for hierarchical multi-label text classification neural network. Expert Systems with Applications, page 115922.
  37. 37.Rui Zhao, Xiao Wei, Cong Ding, and Yongqi Chen. 2021. Hierarchical multi-label text classification: Self-adaption semantic awareness network integrating text topic and label level information. In International Conference on Knowledge Science, Engineering and Management, pages 406–418. Springer.
  38. 38.Jie Zhou, Chunping Ma, Dingkun Long, Guangwei Xu, Ning Ding, Haoyu Zhang, Pengjun Xie, and Gongshen Liu. 2020. Hierarchy-aware global model for hierarchical text classification. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1106–1117, Online. Association for Computational Linguistics.

Citation

MLA
Wang, Z., et al. “Incorporating Hierarchy into Text Encoder: A Contrastive Learning Approach for Hierarchical Text Classification”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 7109–19, https://doi.org/10.18653/v1/2022.acl-long.491.
APA
Wang, Z., (王培懿), P. W., Huang, L., Sun, X., & Wang, H. (2022). Incorporating Hierarchy into Text Encoder: a Contrastive Learning Approach for Hierarchical Text Classification. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7109–7119. https://doi.org/10.18653/v1/2022.acl-long.491
Chicago
Wang, Z., P. W. (王培懿), L. Huang, X. Sun, and H. Wang. 2022. “Incorporating Hierarchy into Text Encoder: A Contrastive Learning Approach for Hierarchical Text Classification”. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 7109–19. https://doi.org/10.18653/v1/2022.acl-long.491.
Harvard
Wang, Z. et al. (2022) “Incorporating Hierarchy into Text Encoder: a Contrastive Learning Approach for Hierarchical Text Classification”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp. 7109–7119. Available at: https://doi.org/10.18653/v1/2022.acl-long.491.
Vancouver
1. Wang Z, (王培懿) PW, Huang L, Sun X, Wang H (2022) Incorporating Hierarchy into Text Encoder: a Contrastive Learning Approach for Hierarchical Text Classification. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, pp 7109–7119

BibTeX

@inproceedings{wang-etal-2022-incorporating,
    title = "Incorporating Hierarchy into Text Encoder: a Contrastive Learning Approach for Hierarchical Text Classification",
    author = "Wang, Zihan  and
      Wang, Peiyi  and
      Huang, Lianzhe  and
      Sun, Xin  and
      Wang, Houfeng",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.491/",
    doi = "10.18653/v1/2022.acl-long.491",
    pages = "7109--7119"
}
Metadata:ACL Anthology

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/