JointCL: A Joint Contrastive Learning Framework for Zero-Shot Stance Detection
Bin LiangQinglin ZhuXiang LiMin YangLin GuiYulan HeRuifeng Xu
Presents a joint contrastive learning framework that uses stance contrastive learning and prototypical graph structures to transfer stance-reasoning capabilities from known targets to unseen ones, achieving state-of-the-art results in zero-shot stance detection.
Modern organizations and public monitoring systems need to automatically identify viewpoints or attitudes expressed in text towards specific topics, propositions, or entities. However, real-world debates continuously produce new subjects that have no historical training data. Existing stance detection methods struggle when faced with these unobserved topics because they cannot reliably generalize what they learned from familiar topics to entirely new ones during deployment.
The main objective of the article is to demonstrate a joint contrastive learning framework, named JointCL, designed to accurately detect stances on previously unseen targets. The article evaluates how combining context-based stance reasoning with target-based relational representations improves performance across zero-shot, few-shot, and cross-target settings.
The evaluated approach uses language representations generated by a pre-trained transformer model and enhances them through two complementary strategies. First, a stance contrastive learning mechanism groups instances sharing the same viewpoint while pushing away differing viewpoints. Second, the system clusters training examples into representative concept prototypes and builds graphs connecting each instance to these prototypes. Using a graph attention network and a novel edge-oriented contrastive objective, the model learns structural relationships between known topics and prototypes and transfers those patterns to unseen targets. The authors validated this framework across three benchmark datasets covering diverse general topics, political debates, and corporate financial mergers.
The experimental findings show that the proposed framework consistently establishes new state-of-the-art benchmarks in zero-shot stance detection. On the large, multi-topic benchmark, the framework achieved an overall Macro F1 score of 72.3%, significantly outperforming established neural and graph baselines. Ablation analyses showed that removing the stance contrastive module or the prototypical graph module caused noticeable performance drops across all datasets, confirming both components are necessary. The framework also generalized effectively to few-shot scenarios with a leading 71.5% F1 score and achieved superior performance across all cross-target evaluation pairs.
These findings imply that automated sentiment and stance analysis systems can reliably handle fast-emerging public topics without requiring expensive, time-consuming data re-labeling or target-specific retraining. The success of bridging known and unknown targets through shared prototype graphs demonstrates that relational structure is critical when contextual text alone is insufficient to deduce a viewpoint. Consequently, adopting this unified framework reduces deployment latency and ongoing annotation costs for monitoring systems.
For practical implementation, organizations deploying stance detection should adopt prototype-driven graph contrastive architectures rather than standard fine-tuning or adversarial methods. Teams should tune the number of prototype clusters to match the diversity of their target domain, using larger cluster counts for broad topic distributions and smaller counts for narrow, domain-specific settings. A sensible next step is to run domain-specific pilot deployments to confirm prototype stability under evolving real-time data streams.
Confidence in these findings is supported by rigorous multi-run evaluations and statistical significance testing across diverse benchmarks. However, practitioners should note that performance remains sensitive to hyper-parameter choices, such as cluster quantity and loss weighting, and zero-shot accuracy is inherently lower than cross-target accuracy where target relationships are partially known in advance.
- Paper: Prototypical Networks for Few-shot Learning, Jake Snell et al. (2017). Its class-prototype construction provides the foundation for understanding JointCL’s use of prototypes to transfer target-based representations to unseen targets.
- Paper: Graph Contrastive Learning with Augmentations, Yuning You et al. (2020). GraphCL introduces contrastive representation learning on graph-structured data, preparing you to follow JointCL’s prototypical graph contrastive component.
- Paper: Supervised Contrastive Learning, Prannay Khosla et al. (2020). Its supervised contrastive objective clarifies the class-aware contrastive learning principles JointCL adapts to improve stance-feature generalization.
- Paper: SimCSE: Simple Contrastive Learning of Sentence Embeddings, Tianyu Gao et al. (2021). SimCSE shows how contrastive objectives can shape transferable sentence representations, a useful prerequisite for understanding JointCL’s stance representation learning.
No sufficiently relevant recommendations were found.
