HPT: Hierarchy-aware Prompt Tuning for Hierarchical Text Classification
Zihan WangPeiyi WangTianyu LiuBinghuai LinYunbo CaoZhifang SuiHoufeng Wang
Proposes a hierarchy-aware prompt tuning framework that reformulates hierarchical text classification into a multi-label masked language modeling task using soft prompts and a zero-bounded cross-entropy loss, significantly improving performance under low-resource and class-imbalanced conditions.
Hierarchical text classification involves assigning documents to multiple categories organized in a structured tree or taxonomy. Traditional approaches often rely on fine-tuning pretrained language models; however, this standard paradigm creates a disconnect between the model's original training objective—predicting masked words in sentences—and the complex, structured classification task. Consequently, existing models struggle to fully exploit language model knowledge, particularly when dealing with deep hierarchies, severely imbalanced classes, or scarce training data.
The main objective of the article is to demonstrate that prompt tuning, when adapted to account for category hierarchies and multi-label outputs, significantly improves hierarchical text classification performance across diverse benchmarks.
To bridge the gap between pretraining and downstream classification, the authors developed Hierarchy-aware Prompt Tuning (HPT). The approach reconfigures the classification task into a masked language modeling format by constructing layer-by-layer virtual prompts and using a graph neural network to inject taxonomy relationships directly into the prompt representations. Additionally, it applies a specialized zero-bounded multi-label loss function to naturally rank target categories above non-target ones. The authors evaluated this framework on three standard benchmark datasets with varying hierarchy depths: Web of Science (2 levels), RCV1-V2 (4 levels), and NYTimes (8 levels), benchmarking it against standard fine-tuning models, advanced graph-based architectures, and basic prompt-tuning methods.
The analysis yielded several key findings. First, HPT achieved new state-of-the-art results across all three benchmarks, delivering its largest gains on the most complex taxonomies, such as improving the Macro-F1 score on the NYTimes dataset from 67.96 to 70.42. Second, standard prompt-tuning baselines without hierarchy awareness also matched or outperformed conventional fine-tuning baselines, confirming that prompt-based formulation better activates pretrained model representations. Third, HPT demonstrated strong robustness in data-constrained scenarios: when evaluated on a low-resource setting using only 10% of the training data, HPT outperformed strong baselines by substantially wider margins (for instance, widening its lead on the RCV1-V2 dataset from 2.13 points in the full-data setting to 6.09 points in the 10% data setting). Finally, ablation analyses showed that injecting structural hierarchy and using the customized ranking loss were crucial for accurately predicting rare, long-tail categories.
These findings indicate that aligning downstream classification tasks closely with the original pretraining mechanics of language models unlocks untapped predictive power without adding model parameters. For organizations managing large-scale document repositories, this approach reduces the risk of misclassifying niche categories and lowers the cost of manual labeling by maintaining high accuracy even with minimal labeled training data.
Organizations seeking to classify structured text taxonomies should consider adopting hierarchy-aware prompt tuning frameworks, particularly for imbalanced datasets or early-stage initiatives with limited training examples. Before wide operational deployment, teams should evaluate trade-offs regarding text length and taxonomy structure. The framework is currently constrained to language models pretrained on masked word prediction and requires dedicated input tokens for each level of the taxonomy, which slightly reduces the maximum length of the input text and may limit applicability on exceptionally deep hierarchies.
- Paper: The Power of Scale for Parameter-Efficient Prompt Tuning, Brian Lester et al. (2021). This study establishes how continuous soft prompts adapt frozen pretrained models, the core parameter-efficient mechanism HPT augments with hierarchy.
- Paper: P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks, Xiao Liu et al. (2022). Its deep prompt-tuning framework clarifies how prompts can be inserted across transformer layers, a design foundation for HPT’s hierarchical soft prompts.
- Paper: Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing, Pengfei Liu et al. (2021). This survey maps prompt-learning methods, including continuous prompts and cloze-style objectives, that HPT combines for hierarchical classification.
- Paper: Exploiting Cloze-Questions for Few-Shot Text Classification and Natural Language Inference, Timo Schick et al. (2020). Pattern-Exploiting Training shows how classification can be recast as masked-language-model cloze prediction, motivating HPT’s multi-label MLM formulation.
- Paper: GPT Understands, Too, Xiao Liu et al. (2021). P-Tuning introduces trainable continuous prompt embeddings, the soft-prompt technique underlying HPT’s hierarchy-aware adaptation.
- Paper: Incorporating Hierarchy into Text Encoder: a Contrastive Learning Approach for Hierarchical Text Classification, Zihan Wang et al. (2022). This hierarchical text-classification method shows how label-taxonomy structure can guide learned text representations, clarifying the hierarchy HPT encodes in prompts.
No sufficiently relevant recommendations were found.
