HPT: Hierarchy-aware Prompt Tuning for Hierarchical Text Classification

Zihan WangPeiyi WangTianyu LiuBinghuai LinYunbo CaoZhifang SuiHoufeng Wang

article2022EMNLP70 citations

Proposes a hierarchy-aware prompt tuning framework that reformulates hierarchical text classification into a multi-label masked language modeling task using soft prompts and a zero-bounded cross-entropy loss, significantly improving performance under low-resource and class-imbalanced conditions.

Listen

Hierarchical text classification involves assigning documents to multiple categories organized in a structured tree or taxonomy. Traditional approaches often rely on fine-tuning pretrained language models; however, this standard paradigm creates a disconnect between the model's original training objective—predicting masked words in sentences—and the complex, structured classification task. Consequently, existing models struggle to fully exploit language model knowledge, particularly when dealing with deep hierarchies, severely imbalanced classes, or scarce training data.

The main objective of the article is to demonstrate that prompt tuning, when adapted to account for category hierarchies and multi-label outputs, significantly improves hierarchical text classification performance across diverse benchmarks.

To bridge the gap between pretraining and downstream classification, the authors developed Hierarchy-aware Prompt Tuning (HPT). The approach reconfigures the classification task into a masked language modeling format by constructing layer-by-layer virtual prompts and using a graph neural network to inject taxonomy relationships directly into the prompt representations. Additionally, it applies a specialized zero-bounded multi-label loss function to naturally rank target categories above non-target ones. The authors evaluated this framework on three standard benchmark datasets with varying hierarchy depths: Web of Science (2 levels), RCV1-V2 (4 levels), and NYTimes (8 levels), benchmarking it against standard fine-tuning models, advanced graph-based architectures, and basic prompt-tuning methods.

The analysis yielded several key findings. First, HPT achieved new state-of-the-art results across all three benchmarks, delivering its largest gains on the most complex taxonomies, such as improving the Macro-F1 score on the NYTimes dataset from 67.96 to 70.42. Second, standard prompt-tuning baselines without hierarchy awareness also matched or outperformed conventional fine-tuning baselines, confirming that prompt-based formulation better activates pretrained model representations. Third, HPT demonstrated strong robustness in data-constrained scenarios: when evaluated on a low-resource setting using only 10% of the training data, HPT outperformed strong baselines by substantially wider margins (for instance, widening its lead on the RCV1-V2 dataset from 2.13 points in the full-data setting to 6.09 points in the 10% data setting). Finally, ablation analyses showed that injecting structural hierarchy and using the customized ranking loss were crucial for accurately predicting rare, long-tail categories.

These findings indicate that aligning downstream classification tasks closely with the original pretraining mechanics of language models unlocks untapped predictive power without adding model parameters. For organizations managing large-scale document repositories, this approach reduces the risk of misclassifying niche categories and lowers the cost of manual labeling by maintaining high accuracy even with minimal labeled training data.

Organizations seeking to classify structured text taxonomies should consider adopting hierarchy-aware prompt tuning frameworks, particularly for imbalanced datasets or early-stage initiatives with limited training examples. Before wide operational deployment, teams should evaluate trade-offs regarding text length and taxonomy structure. The framework is currently constrained to language models pretrained on masked word prediction and requires dedicated input tokens for each level of the taxonomy, which slightly reduces the maximum length of the input text and may limit applicability on exceptionally deep hierarchies.

No sufficiently relevant recommendations were found.

Cover for HPT: Hierarchy-aware Prompt Tuning for Hierarchical Text Classification

Abstract

Hierarchical text classification (HTC) is a challenging subtask of multi-label classification due to its complex label hierarchy. Recently, the pretrained language models (PLM) have been widely adopted in HTC through a fine-tuning paradigm. However, in this paradigm, there exists a huge gap between the classification tasks with sophisticated label hierarchy and the masked language model (MLM) pre-training tasks of PLMs and thus the potential of PLMs cannot be fully tapped. To bridge the gap, in this paper, we propose HPT, a Hierarchy-aware Prompt Tuning method to handle HTC from a multi-label MLM perspective. Specifically, we construct a dynamic virtual template and label words that take the form of soft prompts to fuse the label hierarchy knowledge and introduce a zero-bounded multi-label cross-entropy loss to harmonize the objectives of HTC and MLM. Extensive experiments show HPT achieves state-of-the-art performances on 3 popular HTC datasets and is adept at handling the imbalance and low resource situations. Our code is available at https://github.com/wzh9969/HPT.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 2.1 Hierarchical Text Classification
  • 2.2 Prompt tuning
  • 3 Preliminaries
  • 3.1 Problem Definition
  • 3.2 Vanilla Fine Tuning for HTC
  • 3.3 Prompt Tuning for HTC
  • 4 Methodology
  • 4.1 Hierarchy-aware Prompt
  • 4.1.1 Hierarchy Constraint
  • 4.1.2 Hierarchy Injection
  • 4.2 Zero-bounded Multi-label Cross-entropy Loss
  • 5 Experiments
  • 5.1 Experiment Setup
  • 5.2 Main Results
  • 5.3 Ablation Study
  • 5.4 Interpreting on Representation Space
  • 5.5 Results on Imbalanced Hierarchy
  • 5.6 Results on Low Resource Setting
  • 6 Conclusion
  • Limitations
  • Acknowledgements
  • References
  • A Data Statistics
  • B Example of Different Prompt Methods
  • C Discussion on Different Connections of Hierarchy Injection
  • D Ablation results on Web of Science and RCV1-V2

Knowls

  1. Knowl 1 — HPT reframes tree-structured classification as hierarchy-aware multi-label MLM

    model/method

    HPT addresses hierarchical text classification (HTC) on a rooted, tree-like label hierarchy: every non-root label has exactly one parent, and a text may receive labels forming one or more paths in the hierarchy. Instead of treating the labels as independent binary decisions, HPT adapts a masked language model (MLM) to make multi-label predictions at hierarchy levels. Its main components are a depth-dependent soft prompt, label-hierarchy information injected into that prompt, and a multi-label loss compatible with MLM-style prediction.

  2. Knowl 2 — Depth-specific prediction slots constrain the prompt to the label hierarchy

    model/method

    For a label tree with LL levels and input text xx, HPT appends a learned template vector and a prediction token for every level, giving the pattern [CLS] x [SEP] [V1] [PRED]⋯[VL] [PRED] [SEP][\mathrm{CLS}]\ x\ [\mathrm{SEP}]\ [V_1]\ [\mathrm{PRED}]\cdots[V_L]\ [\mathrm{PRED}]\ [\mathrm{SEP}]. Thus, unlike a soft prompt with a fixed number of template vectors, the number of HPT template vectors is determined by the hierarchy depth. The embedding of [PRED][\mathrm{PRED}] is initialized from BERT's [MASK][\mathrm{MASK}] embedding. Each label yy has a learned virtual label-word embedding vyv_y, initialized as the average of the embeddings of the tokens in that label's name. The prediction slot for level mm is verbalized only by labels at level mm; labels from other levels are excluded from that slot. This separates label prediction by depth while retaining multi-label prediction within each level.

  3. Knowl 3 — A graph attention network injects label connectivity into each level's prompt

    model/method

    To represent relationships among labels as well as their depths, HPT builds a graph from the label hierarchy and adds one virtual node for each hierarchy level. Virtual node tmt_m is connected to every label node at level mm. Label nodes use their virtual label-word embeddings as input features, while each virtual level node uses its learned template embedding. A stacked graph attention network (GAT) propagates information over this augmented graph; the resulting representation of each virtual node is added residually to its original template embedding. The updated template vectors are supplied to BERT. The paper's experiments use a single GAT layer. This mechanism lets a level's prompt incorporate information from connected labels rather than representing the hierarchy only through separate depth-specific slots.

  4. Knowl 4 — Zero-bounded multi-label cross-entropy separates positive and negative label scores

    equation

    For hierarchy level mm, let Nm+N_m^+ and Nm−N_m^- be the target and non-target labels for an input, respectively. Let hPmh_P^m be BERT's hidden vector at that level's prediction token, and let vyv_y and bm,yb_{m,y} be the learned label embedding and bias for label yy. The score is sm,y=vy⊤hPm+bm,ys_{m,y}=v_y^\top h_P^m+b_{m,y}. HPT's zero-bounded multi-label cross-entropy (ZMLCE) loss at that level is

    LZMLCEm=log⁡(1+∑y∈Nm−exp⁡(sm,y))+log⁡(1+∑y∈Nm+exp⁡(−sm,y)).\mathcal{L}_{\mathrm{ZMLCE}}^m =\log\left(1+\sum_{y\in N_m^-}\exp(s_{m,y})\right) +\log\left(1+\sum_{y\in N_m^+}\exp(-s_{m,y})\right).

    The loss encourages non-target scores to fall below the fixed anchor score zero and target scores to rise above zero. Unlike a loss requiring a known number of positive labels at inference, this threshold-based formulation permits HPT to predict each label with the rule sm,y>0s_{m,y}>0.

  5. Knowl 5 — HPT jointly trains hierarchical prediction and masked-word reconstruction

    model/method

    HPT's training objective sums the ZMLCE losses over the LL hierarchy levels and adds BERT's masked language modeling loss: Lall=∑m=1LLZMLCEm+LMLM\mathcal{L}_{\mathrm{all}}=\sum_{m=1}^{L}\mathcal{L}_{\mathrm{ZMLCE}}^m+\mathcal{L}_{\mathrm{MLM}}. The MLM loss is computed by randomly masking 15% of the input text's words. The added MLM objective retains the pretrained masked-word recovery task while the level-specific ZMLCE terms train the hierarchy-aware multi-label predictions.

  6. Knowl 6 — HPT achieves the strongest reported test F1 across three HTC datasets

    data/table

    The test-set comparison reports Micro-F1 and Macro-F1 for Web-of-Science (WOS, hierarchy depth 2), RCV1-V2 (depth 4), and NYTimes (NYT, depth 8). HPT has the highest score in each of the six dataset–metric comparisons, including the largest margins on the deepest hierarchy, NYT. The reported results use BERT-base-uncased; HPT uses one GAT layer, batch size 16, and Adam with learning rate 3×10−53\times10^{-5}. Models are evaluated on the development set after each epoch and training stops after six epochs without an increase in development Macro-F1. The prompt and other hyperparameters were not tuned.

    WOS RCV1-V2 NYT
    Model Micro-F1 Macro-F1 Micro-F1 Macro-F1 Micro-F1 Macro-F1
    TextRCNN 83.55 76.99 81.57 59.25 70.83 56.18
    HiAGM 85.82 80.28 83.96 63.35 74.97 60.83
    HTCInfoMax 85.58 80.05 83.51 62.71 – –
    HiMatch 86.20 80.53 84.73 64.11 – –
    BERT 85.63 79.07 85.65 67.02 78.24 66.08
    BERT+HiAGM 86.04 80.19 85.58 67.93 78.64 66.76
    BERT+HTCInfoMax 86.30 79.97 85.53 67.09 78.75 67.31
    BERT+HiMatch 86.70 81.06 86.33 68.66 – –
    HGCLR 87.11 81.20 86.49 68.31 78.86 67.96
    BERT+HardPrompt 86.39 80.43 86.78 68.78 79.45 67.99
    BERT+SoftPrompt 86.57 80.75 86.53 68.34 78.95 68.21
    HPT 87.16 81.93 87.26 69.53 80.42 70.42
  7. Knowl 7 — Ablations support the value of hierarchy injection and ZMLCE

    data/table

    These development-set ablations remove one HPT component, replace ZMLCE with binary cross-entropy (BCE), or add random label-to-level connections. Removing hierarchy injection or replacing ZMLCE with BCE lowers Macro-F1 on all three datasets. On NYT, removing hierarchy injection lowers Macro-F1 from 71.07 to 69.71; random connections lower it to 69.42, worse than removing hierarchy injection altogether. The authors interpret this as evidence that structurally inconsistent connections can be harmful, while the hierarchy-aware connections provide useful information. The values below are Micro-F1 / Macro-F1.

    Dataset Variant Micro-F1 Macro-F1
    NYT HPT 80.49 71.07
    NYT Remove hierarchy constraint 80.32 70.58
    NYT Remove hierarchy injection 80.41 69.71
    NYT Replace ZMLCE with BCE 79.74 70.40
    NYT Remove MLM loss 80.16 70.78
    NYT Random connections 80.12 69.42
    WOS HPT 87.88 81.68
    WOS Remove hierarchy constraint 87.34 81.27
    WOS Remove hierarchy injection 87.58 81.54
    WOS Replace ZMLCE with BCE 87.17 80.78
    WOS Remove MLM loss 87.22 81.36
    WOS Random connections 87.56 81.42
    RCV1-V2 HPT 88.37 70.12
    RCV1-V2 Remove hierarchy constraint 87.62 69.04
    RCV1-V2 Remove hierarchy injection 87.57 68.53
    RCV1-V2 Replace ZMLCE with BCE 87.79 68.12
    RCV1-V2 Remove MLM loss 87.83 69.76
    RCV1-V2 Random connections 88.22 68.86
  8. Knowl 8 — HPT improves performance on sparse labels and imbalanced hierarchy levels

    empirical result

    On the NYT development set, the authors analyze label groups by hierarchy depth and by the number of training examples per label. The middle levels, depths 3 and 4, contain more labels than the shallow or deepest levels, and all compared models perform poorly on these groups; HPT chiefly improves the middle-level results. Grouping labels by training frequency also shows larger improvements for labels with fewer training examples, consistent with alleviating the long-tail effect. The NYT label with the most training examples has over 100 times as many examples as the least frequent label. The paper reports these trends graphically rather than giving exact per-group F1 values.

  9. Knowl 9 — HPT retains an advantage when trained on only 10% of the data

    empirical result

    For a low-resource evaluation, the authors randomly use 10% of the training data for each of WOS, RCV1-V2, and NYT, leaving the other experimental settings unchanged. Scores are averaged over three runs, with standard deviations reported. HPT outperforms the comparison models on all three datasets and is reported to have lower run-to-run standard deviation. The authors also report that, on RCV1-V2, HPT's Macro-F1 advantage over BERT+HTCInfoMax grows from 2.13 points in the full-resource setting to 6.09 points in the low-resource setting. The paper presents the low-resource scores graphically without listing their exact values.

  10. Knowl 10 — HPT depends on MLM backbones and consumes input length for prompts

    limitation

    HPT relies on a pretrained language model with an MLM objective, so it is not directly applicable to pretrained models that do not incorporate MLM. Its additional template tokens also reduce the portion of the maximum sequence length available to the input text. Because the number of template vectors grows with hierarchy depth, the authors caution that HPT may fail on datasets with extreme hierarchy depth.

Coverage note — The paper's nearest-neighbor label-embedding interpretation and alternative depth-increasing virtual-node connections are omitted: they are supporting analyses rather than load-bearing method or benchmark results.

References

  1. 1.Siddhartha Banerjee, Cem Akkaya, Francisco Perez-Sorrosal, and Kostas Tsioutsiouliklis. 2019. Hierarchical transfer learning for multi-label text classification. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 6295–6300, Florence, Italy. Association for Computational Linguistics.
  2. 2.Boli Chen, Xin Huang, Lin Xiao, Zixin Cai, and Liping Jing. 2020. Hyperbolic interaction model for hierarchical multi-label classification. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pages 7496–7503.
  3. 3.Haibin Chen, Qianli Ma, Zhenxi Lin, and Jiangyue Yan. 2021. Hierarchy-aware label semantics matching network for hierarchical text classification. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4370–4379, Online. Association for Computational Linguistics.
  4. 4.Xiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng, Yunzhi Yao, Chuanqi Tan, Fei Huang, Luo Si, and Huajun Chen. 2022. Knowprompt: Knowledge-aware prompt-tuning with synergistic optimization for relation extraction. In Proceedings of the ACM Web Conference 2022, pages 2778–2788.
  5. 5.Zhongfen Deng, Hao Peng, Dongxiao He, Jianxin Li, and Philip Yu. 2021. HTCInfoMax: A global model for hierarchical text classification via information maximization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3259–3265, Online. Association for Computational Linguistics.
  6. 6.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  7. 7.Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. Making pre-trained language models better few-shot learners. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3816–3830, Online. Association for Computational Linguistics.
  8. 8.Siddharth Gopal and Yiming Yang. 2013. Recursive regularization for large-scale classification with hierarchical and graphical dependencies. In Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining, pages 257–265.
  9. 9.Karen Hambardzumyan, Hrant Khachatrian, and Jonathan May. 2021. WARP: Word-level Adversarial ReProgramming. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4921–4933, Online. Association for Computational Linguistics.
  10. 10.Rie Johnson and Tong Zhang. 2015. Effective use of word order for text categorization with convolutional neural networks. In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 103–112, Denver, Colorado. Association for Computational Linguistics.
  11. 11.Thomas N. Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In 5th International Conference on Learning Representations (ICLR).
  12. 12.Kamran Kowsari, Donald E Brown, Mojtaba Heidarysafa, Kiana Jafari Meimandi, Matthew S Gerber, and Laura E Barnes. 2017. Hdltex: Hierarchical deep learning for text classification. In 2017 16th IEEE international conference on machine learning and applications (ICMLA), pages 364–371. IEEE.
  13. 13.Siwei Lai, Liheng Xu, Kang Liu, and Jun Zhao. 2015. Recurrent convolutional neural networks for text classification. In Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, pages 2267–2273.
  14. 14.David D Lewis, Yiming Yang, Tony Russell-Rose, and Fan Li. 2004. Rcv1: A new benchmark collection for text categorization research. Journal of machine learning research, 5(Apr):361–397.
  15. 15.Yuning Mao, Jingjing Tian, Jiawei Han, and Xiang Ren. 2019. Hierarchical text classification with reinforced label assignment. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 445–455, Hong Kong, China. Association for Computational Linguistics.
  16. 16.Guanghui Qin and Jason Eisner. 2021. Learning how to ask: Querying lms with mixtures of soft prompts. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 5203–5212.
  17. 17.Evan Sandhaus. 2008. The new york times annotated corpus. Linguistic Data Consortium, Philadelphia, 6(12):e26752.
  18. 18.Timo Schick and Hinrich Schütze. 2021. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 255–269.
  19. 19.Kazuya Shimura, Jiyi Li, and Fumiyo Fukumoto. 2018. HFT-CNN: Learning hierarchical category structure for multi-label short text categorization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 811–816, Brussels, Belgium. Association for Computational Linguistics.
  20. 20.Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4222–4235, Online. Association for Computational Linguistics.
  21. 21.Carlos N Silla and Alex A Freitas. 2011. A survey of hierarchical classification across different application domains. Data Mining and Knowledge Discovery, 22(1):31–72.
  22. 22.Jianlin Su. 2020. Extending “softmax+cross entropy” to multi-label classification problem. https://spaces.ac.cn/archives/7359.
  23. 23.Yifan Sun, Changmao Cheng, Yuhan Zhang, Chi Zhang, Liang Zheng, Zhongdao Wang, and Yichen Wei. 2020. Circle loss: A unified perspective of pair similarity optimization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6398–6407.
  24. 24.Zihan Wang, Peiyi Wang, Lianzhe Huang, Xin Sun, and Houfeng Wang. 2022. Incorporating hierarchy into text encoder: a contrastive learning approach for hierarchical text classification. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7109–7119, Dublin, Ireland. Association for Computational Linguistics.
  25. 25.Jonatas Wehrmann, Ricardo Cerri, and Rodrigo Barros. 2018. Hierarchical multi-label classification networks. In International Conference on Machine Learning, pages 5075–5084. PMLR.
  26. 26.Jiawei Wu, Wenhan Xiong, and William Yang Wang. 2019. Learning to learn and predict: A meta-learning approach for multi-label classification. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4354–4364, Hong Kong, China. Association for Computational Linguistics.
  27. 27.Xinyi Zhang, Jiahao Xu, Charlie Soh, and Lihui Chen. 2021. La-hcn: Label-based attention for hierarchical multi-label text classification neural network. Expert Systems with Applications, page 115922.
  28. 28.Rui Zhao, Xiao Wei, Cong Ding, and Yongqi Chen. 2021. Hierarchical multi-label text classification: Self-adaption semantic awareness network integrating text topic and label level information. In International Conference on Knowledge Science, Engineering and Management, pages 406–418. Springer.
  29. 29.Jie Zhou, Chunping Ma, Dingkun Long, Guangwei Xu, Ning Ding, Haoyu Zhang, Pengjun Xie, and Gongshen Liu. 2020. Hierarchy-aware global model for hierarchical text classification. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1106–1117, Online. Association for Computational Linguistics.

Citation

MLA
Wang, Z., et al. “HPT: Hierarchy-aware Prompt Tuning for Hierarchical Text Classification”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 3740–51, https://doi.org/10.18653/v1/2022.emnlp-main.246.
APA
Wang, Z., (王培懿), P. W., Liu, T., Lin, B., Cao, Y., Sui, Z., & Wang, H. (2022). HPT: Hierarchy-aware Prompt Tuning for Hierarchical Text Classification. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3740–3751. https://doi.org/10.18653/v1/2022.emnlp-main.246
Chicago
Wang, Z., P. W. (王培懿), T. Liu, et al. 2022. “HPT: Hierarchy-aware Prompt Tuning for Hierarchical Text Classification”. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 3740–51. https://doi.org/10.18653/v1/2022.emnlp-main.246.
Harvard
Wang, Z. et al. (2022) “HPT: Hierarchy-aware Prompt Tuning for Hierarchical Text Classification”, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp. 3740–3751. Available at: https://doi.org/10.18653/v1/2022.emnlp-main.246.
Vancouver
1. Wang Z, (王培懿) PW, Liu T, Lin B, Cao Y, Sui Z, Wang H (2022) HPT: Hierarchy-aware Prompt Tuning for Hierarchical Text Classification. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, pp 3740–3751

BibTeX

@inproceedings{wang-etal-2022-hpt,
    title = "{HPT}: Hierarchy-aware Prompt Tuning for Hierarchical Text Classification",
    author = "Wang, Zihan  and
      Wang, Peiyi  and
      Liu, Tianyu  and
      Lin, Binghuai  and
      Cao, Yunbo  and
      Sui, Zhifang  and
      Wang, Houfeng",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.246/",
    doi = "10.18653/v1/2022.emnlp-main.246",
    pages = "3740--3751"
}
Metadata:ACL Anthology

Source Code

This paper has an official code repository available. Click below to access the source code.

View Repository

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/