Zero-Shot Rumor Detection with Propagation Structure via Prompt Learning

Hongzhan LinPengyao YiJing MaHaiyun JiangZiyang LuoShuming ShiRuifang Liu

article2023AAAI83 citations

Proposes a response-aware prompt learning framework that integrates domain-invariant propagation structures and hierarchical prompt encoding into multilingual pre-trained language models to accurately detect rumors across unseen languages and domains without target-specific annotations.

Listen

During fast-moving global crises and breaking events, rumors spread rapidly across social media platforms in multiple languages, threatening public health, safety, and institutional trust. Existing automated rumor detection systems rely heavily on large volumes of annotated data, which are typically concentrated in high-resource languages like English and Chinese. As a result, standard tools fail when confronted with unforeseen events emerging in low-resource or minority languages where labeled training examples do not exist, and human fact-checkers cannot scale quickly enough to assess the claims.

The article demonstrates an automated zero-shot rumor detection framework that identifies unverified claims across new languages and domains without requiring target-language training annotations. It evaluates how prompt-based learning and the structural dynamics of social media conversations can overcome the scarcity of labeled data during breaking events.

To achieve this, the researchers developed the Response-aware Prompt Learning framework. The approach decouples language-dependent grammar from shared semantic meaning by freezing the lower layers of a multilingual language model and fine-tuning the upper layers. It structures user replies using tree-based search patterns, models both absolute and relative propagation positions within conversation threads, applies adversarial data augmentation to handle social media noise, and uses class prototypes to classify rumors without language-specific label tuning. The framework was evaluated across four public datasets and a newly constructed low-resource benchmark containing Cantonese and Arabic social media threads related to COVID-19.

The framework consistently outperformed all conventional transfer learning, adapter, and baseline prompt-tuning methods across multiple evaluation metrics. Among structural ranking strategies, breadth-first traversal yielded the most robust and accurate results, outperforming depth-first and purely chronological orderings. Crucially, the system demonstrated strong early-detection capabilities, reaching saturated predictive accuracy with only 20 reply posts or within 4 hours of an initial post, significantly faster than baseline methods.

These findings indicate that user response patterns and conversational structures carry universal, language-agnostic signals that can be leveraged to detect misinformation. For platform operators, policymakers, and public health organizations, this provides a cost-effective and scalable pathway to monitor emergent crises globally without maintaining expensive fact-checking or annotation pipelines for every regional dialect. Early automated flagging also reduces response timelines, allowing timely interventions before false narratives become deeply entrenched.

Organizations monitoring misinformation should incorporate conversation propagation structures and hierarchical prompt architectures into their content moderation workflows. As a next step, deploying a pilot system to evaluate live streaming feeds during unfolding events would help determine practical operational trade-offs between speed and accuracy.

The approach is bounded by the standard context length limits of underlying language models and relies on the existence of initial user interactions. However, across the evaluated benchmarks, the framework delivers strong, consistent performance and offers a reliable foundation for zero-shot rumor detection.

Cover for Zero-Shot Rumor Detection with Propagation Structure via Prompt Learning

Abstract

The spread of rumors along with breaking events seriously hinders the truth in the era of social media. Previous studies reveal that due to the lack of annotated resources, rumors presented in minority languages are hard to be detected. Furthermore, the unforeseen breaking events not involved in yesterday's news exacerbate the scarcity of data resources. In this work, we propose a novel zero-shot framework based on prompt learning to detect rumors falling in different domains or presented in different languages. More specifically, we firstly represent rumor circulated on social media as diverse propagation threads, then design a hierarchical prompt encoding mechanism to learn language-agnostic contextual representations for both prompts and rumor data. To further enhance domain adaptation, we model the domain-invariant structural features from the propagation threads, to incorporate structural position representations of influential community response. In addition, a new virtual response augmentation method is used to improve model training. Extensive experiments conducted on three real-world datasets demonstrate that our proposed model achieves much better performance than state-of-the-art methods and exhibits a superior capacity for detecting rumors at early stages.

Table of Contents

  • 1 Introduction
  • 2 Related Work
  • 3 Problem Statement and Background
  • 4 Our Approach
  • 4.1 Response Ranking
  • 4.2 Hierarchical Prompt Encoding
  • 4.3 Propagation Position Modeling
  • 4.4 Response Augmentation
  • 4.5 Model Training
  • 5 Experiments
  • 5.1 Datasets
  • 5.2 Experimental Setup
  • 5.3 Rumor Detection Performance
  • 5.4 Ablation Study
  • 5.5 Early Detection
  • 5.6 Discussion
  • 6 Conclusion and Future Work
  • Acknowledgments
  • References

Knowls

  1. Knowl 1 — Zero-shot rumor detection task and prompt formulation

    definition

    The paper defines zero-shot rumor detection as transferring rumor knowledge from a labeled source dataset to an unlabeled target dataset whose language and domain differ from the source. The source dataset is Ds={C1s,…,CMs}D_s=\{C_1^s,\ldots,C_M^s\}, where each event is Cis=(yi,ci,T(ci))C_i^s=(y_i,c_i,T(c_i)): yi∈{rumor,non-rumor}y_i\in\{\text{rumor},\text{non-rumor}\} is the veracity label, cic_i is the claim, and T(ci)=[x1s,…,xms]T(c_i)=[x_1^s,\ldots,x_m^s] is its chronologically ordered sequence of responsive microblog posts. The target dataset is Dt={C1t,…,CNt}D_t=\{C_1^t,\ldots,C_N^t\}, where each target event contains a claim and response thread but no training-time label.

    The method converts classification into masked-language-model prediction. A language-independent template pp, such as “For this [MASK] story,” is spliced with a claim and its responses to form an input c^\hat c. Let VV be a set of label words, Vy⊆VV_y\subseteq V the words associated with veracity class yy, and gg a verbalizer that maps masked-token probabilities to class probabilities. The prompt formulation is

    P(y∣c^)=g(P([MASK]=v∣c^)∣v∈Vy).P(y\mid\hat c)=g\left(P([\mathrm{MASK}]=v\mid\hat c)\mid v\in V_y\right).

    The proposed system replaces manually selected language-specific label words with learned class prototypes, but retains the masked-language-model formulation for zero-shot inference.

  2. Knowl 2 — Response-aware Prompt Learning framework

    model/method

    The proposed Response-aware Prompt Learning (RPL) framework detects a claim by jointly modeling its text, ranked community responses, multilingual prompt representations, and propagation structure. The architecture illustrated in Figure 2 on page 3 has four connected components: (1) response ranking selects a propagation thread from the available responses; (2) Hierarchical Prompt Encoding uses a frozen lower portion and trainable upper portion of a multilingual language model to separate syntactic and semantic transfer; (3) propagation-position modeling injects absolute tree depth and local relative relationships into the trainable semantic encoder; and (4) Virtual Response Augmentation creates adversarially perturbed response representations during training.

    For source events, the model obtains a representation of the masked token from the prompt–event interaction and trains it against rumor and non-rumor prototypes. For target events, it compares the masked-token state with the learned prototypes and assigns the closer veracity class. The framework is designed to transfer both language-agnostic claim semantics and domain-invariant interaction patterns from the source language/domain to the target language/domain without target annotations.

  3. Knowl 3 — Diverse propagation-thread response ranking

    algorithm

    RPL represents each claim together with a selected ordering of its responsive posts rather than treating the claim as an isolated text. Given a response thread T(c)T(c), four orderings are considered:

    • Chronological order, which places earlier responses first: [x1,x2,…,xm][x_1,x_2,\ldots,x_m].
    • Inverted chronological order, which places later responses first: [xm,xm−1,…,x1][x_m,x_{m-1},\ldots,x_1].
    • Depth-first order on the propagation tree, which follows information flow from an ancestor to its descendants before moving to another branch.
    • Breadth-first order on the propagation tree, which visits posts level by level and emphasizes interactions among users responding within the same subtree.

    For a propagation tree T(c)=⟨G,E⃗⟩T(c)=\langle G,\vec E\rangle, GG is the set of response-post nodes and E⃗\vec E is the set of directed parent–child propagation relations. The selected sequence is concatenated with the claim and truncated to the maximum input length of the multilingual language model, thereby retaining as many contextually coherent responses as possible. In the paper’s tree example, depth-first ranking yields [x1,x2,x5,x3,x4,x6][x_1,x_2,x_5,x_3,x_4,x_6], whereas breadth-first ranking yields [x1,x3,x2,x4,x5,x6][x_1,x_3,x_2,x_4,x_5,x_6]. The four variants are evaluated as RPL-Cho, RPL-Inv, RPL-Dep, and RPL-Bre; breadth-first ranking is generally the most stable.

  4. Knowl 4 — Hierarchical Prompt Encoding for language-independent transfer

    model/method

    Hierarchical Prompt Encoding (HPE) uses different parts of a multilingual pre-trained language model for syntactic and semantic processing. The paper assumes that lower transformer layers mainly encode language-dependent syntactic information while upper layers contain more transferable semantic information. The lower kk layers are copied from the multilingual language model and frozen. With a template pp, claim cc, and ranked response sequence T(c)T(c), the frozen SynEncoder produces

    Xp=SynEncoder(p),X_p=\mathrm{SynEncoder}(p),

    where Xp∈R∣p∣×dX_p\in\mathbb{R}^{|p|\times d}, ∣p∣|p| is the template length, and dd is the hidden-state dimension, and

    Xcr=SynEncoder([c,T(c)]),X_{cr}=\mathrm{SynEncoder}([c,T(c)]),

    where Xcr∈Ro×dX_{cr}\in\mathbb{R}^{o\times d} and oo is the model’s maximum input length. The template and claim–response event are then concatenated and passed to a trainable semantic encoder initialized from layers k+1k+1 through the top layer of the multilingual model:

    H=SemEncoder([Xp,Xcr]).H=\mathrm{SemEncoder}([X_p,X_{cr}]).

    Here HH contains the contextual states used for masked-token classification. This design maps a simple shared template into a multilingual embedding space before allowing the trainable upper layers to model its semantic interaction with an event, avoiding manually designed rumor-specific templates in every target language. The experiments set k=6k=6.

  5. Knowl 5 — Absolute and relative propagation-position modeling

    model/method

    RPL injects propagation-tree structure into the trainable semantic encoder through two complementary position representations. For a token qq belonging to response post xix_i, the absolute propagation position is the distance from that post to the claim root cc in the propagation tree T(c)T(c):

    abspro(q)=distanceT(c)(xi,c).\mathrm{abspro}(q)=\mathrm{distance}_{T(c)}(x_i,c).

    All tokens in the same response post share the post’s absolute position. The corresponding learnable absolute-position embedding is added to the token representation before semantic encoding.

    Relative propagation position models the local structure around a response post. For a current post, the model distinguishes five relations: an earlier parent, a later child, an earlier sibling, a later sibling, and the post itself. The earlier/later distinction is defined by the ordering within the local subtree. These pairwise relations are incorporated into self-attention using relative-position representations, allowing the semantic encoder to model how users in the same propagation subtree respond to and cross-check a common topic. Absolute depth captures a post’s global location, while relative relations capture local opinion interactions.

  6. Knowl 6 — Virtual Response Augmentation

    algorithm

    Virtual Response Augmentation (ViRA) creates an adversarially perturbed version of each claim–response event to make RPL less sensitive to noisy responses and input-length constraints. Its input is a source training event, the frozen SynEncoder, the current trainable SemEncoder, and the current training loss; its output is a pseudo-augmented response representation and an additional training loss.

    Input: claim–response event, frozen SynEncoder, trainable SemEncoder, current loss
    Output: pseudo-augmented response representation and augmented loss
    Encode the template and the claim–response event with the frozen SynEncoder.
    Apply layer normalization to the resulting embeddings.
    Mask the template and claim positions so that only response-token embeddings can be perturbed.
    Compute the gradient of the current loss with respect to the normalized response embeddings.
    Normalize the gradient to obtain an adversarial perturbation direction.
    Add the scaled normalized perturbation to the normalized response embeddings.
    Keep the template and claim embeddings unchanged.
    Feed the perturbed event and template through the trainable SemEncoder.
    Compute the prototypical-contrastive loss on the perturbed event.
    Return the perturbed event representation and its loss.

    Layer normalization is applied before perturbation because embedding norms vary across data and model sizes. The perturbation is restricted to responsive posts, so ViRA augments the social-context signal without altering the claim or the template.

  7. Knowl 7 — Prototypical-contrastive training objective

    equation

    For a source training event CiC_i, let HimH_i^m be the contextual representation of its [MASK] token after the semantic encoder, yiy_i its ground-truth class, lyl_y the learnable prototype vector for class yy, and S(⋅,⋅)S(\cdot,\cdot) the normalized cosine similarity. The prototypical verbalizer trains the masked-token representation to approach the prototype of its class:

    Lproto=−log⁡exp⁡S(Him,lyi)∑y′exp⁡S(Him,ly′).\mathcal{L}_{\mathrm{proto}}=-\log\frac{\exp S(H_i^m,l_{y_i})}{\sum_{y'}\exp S(H_i^m,l_{y'})}.

    For a minibatch, the contrastive objective pulls representations with the same label together and separates them from representations with different labels:

    Lcon=−1Byi−1∑j1[i≠j]1[yi=yj]log⁡exp⁡S(Him,Hjm)∑j′1[i≠j′]exp⁡S(Him,Hj′m),\mathcal{L}_{\mathrm{con}}=-\frac{1}{B_{y_i}-1}\sum_j \mathbf{1}[i\ne j]\mathbf{1}[y_i=y_j] \log\frac{\exp S(H_i^m,H_j^m)}{\sum_{j'}\mathbf{1}[i\ne j']\exp S(H_i^m,H_{j'}^m)},

    where ByiB_{y_i} is the number of batch examples with label yiy_i and 1[⋅]\mathbf{1}[\cdot] is the indicator function. The source-event loss is

    L=αLproto+(1−α)Lcon,\mathcal{L}=\alpha\mathcal{L}_{\mathrm{proto}}+(1-\alpha)\mathcal{L}_{\mathrm{con}},

    with α=0.5\alpha=0.5 in the experiments. ViRA produces a second event loss L~\widetilde{\mathcal{L}}, and optimization uses Lavg=mean(L+L~)\mathcal{L}_{\mathrm{avg}}=\mathrm{mean}(\mathcal{L}+\widetilde{\mathcal{L}}). The model is trained with AdamW, an initial learning rate of 10−510^{-5}, k=6k=6 frozen SynEncoder layers, and early stopping. At inference, the masked-token state is classified by its similarity to the learned rumor and non-rumor prototypes rather than by manually chosen label words.

  8. Knowl 8 — Datasets and zero-shot experimental protocol

    experimental setup

    The experiments use four existing propagation-thread rumor datasets and one newly constructed low-resource dataset. TWITTER and Twitter-COVID19 contain English claims and tweet conversation threads; WEIBO and Weibo-COVID19 contain Chinese claims and analogous response structures. The authors construct CatAr-COVID19 by taking multilingual Cantonese and Arabic COVID-19 claim collections that lacked propagation threads, collecting their Twitter response threads through the Twitter Academic API, and assigning the original event labels to the collected claim tweets.

    The challenging zero-shot protocol trains on the well-resourced TWITTER or WEIBO dataset and tests without target annotations on Weibo-COVID19, Twitter-COVID19, or CatAr-COVID19. Evaluation uses accuracy, macro-averaged F1, rumor-class F1, and non-rumor-class F1. The comparison includes task-specific fine-tuning, translation-based fine-tuning, supervised-contrastive fine-tuning, adapters, parallel adapters, source-language prompting, translated prompting, and soft prompting. Early-detection comparisons use the same multilingual language model for the baselines and evaluate only responses available by a specified reply-count or elapsed-time checkpoint.

  9. Knowl 9 — Zero-shot rumor-detection performance

    data/table

    The comparative results reported on page 6 evaluate four source–target transfers. Each cell group contains accuracy, macro-F1, rumor F1, and non-rumor F1. The RPL variants consistently outperform the non-RPL baselines in overall transfer performance; the breadth-first variant is especially strong, while inverted chronological ranking is often better than chronological ranking. The results also show that using propagation responses and their structure is beneficial when both the language and domain change.

    TWITTERWeibo-COVID19 TWITTERCatAr-COVID19 WEIBOTwitter-COVID19 WEIBOCatAr-COVID19
    Model Acc. Mac-F1 R-F1 NR-F1 Acc. Mac-F1 R-F1 NR-F1 Acc. Mac-F1 R-F1 NR-F1 Acc. Mac-F1 R-F1 NR-F1
    Vanilla-Ft 0.623 0.585 0.711 0.459 0.518 0.402 0.583 0.220 0.603 0.602 0.619 0.585 0.481 0.481 0.479 0.474
    Translate-Ft 0.639 0.567 0.745 0.388 0.523 0.457 0.637 0.277 0.634 0.574 0.653 0.495 0.505 0.512 0.528 0.496
    Contrast-Ft 0.656 0.582 0.759 0.405 0.584 0.458 0.720 0.196 0.653 0.644 0.699 0.590 0.562 0.561 0.571 0.551
    Adapter 0.644 0.600 0.737 0.463 0.558 0.438 0.665 0.211 0.652 0.612 0.736 0.487 0.548 0.556 0.605 0.508
    Parallel-Adpt 0.651 0.598 0.730 0.467 0.567 0.450 0.701 0.198 0.667 0.653 0.731 0.574 0.579 0.585 0.636 0.534
    Source-Ppt 0.664 0.648 0.722 0.574 0.589 0.564 0.460 0.669 0.670 0.616 0.760 0.472 0.599 0.565 0.688 0.441
    Translate-Ppt 0.650 0.489 0.776 0.201 0.573 0.568 0.519 0.617 0.674 0.651 0.740 0.562 0.604 0.542 0.374 0.711
    Soft-Ppt 0.652 0.574 0.756 0.392 0.590 0.565 0.446 0.683 0.685 0.652 0.758 0.546 0.609 0.575 0.518 0.633
    RPL-Cho 0.713 0.675 0.786 0.563 0.613 0.581 0.455 0.707 0.715 0.689 0.778 0.601 0.634 0.633 0.616 0.650
    RPL-Inv 0.728 0.666 0.810 0.521 0.601 0.592 0.473 0.711 0.733 0.710 0.788 0.632 0.647 0.640 0.586 0.693
    RPL-Dep 0.732 0.689 0.805 0.574 0.640 0.619 0.530 0.708 0.723 0.711 0.771 0.650 0.657 0.636 0.547 0.724
    RPL-Bre 0.745 0.719 0.804 0.634 0.631 0.617 0.544 0.689 0.727 0.697 0.793 0.601 0.672 0.664 0.614 0.714

    For example, with TWITTER as source and Weibo-COVID19 as target, RPL-Bre reaches 0.745 accuracy and 0.719 macro-F1, compared with 0.652 and 0.574 for Soft-Ppt. With WEIBO as source and CatAr-COVID19 as target, RPL-Bre reaches 0.672 accuracy and 0.664 macro-F1, compared with 0.609 and 0.575 for Soft-Ppt.

  10. Knowl 10 — Component ablations and early-detection behavior

    empirical result

    The ablation results on CatAr-COVID19, reported on page 7, use TWITTER or WEIBO as the source and measure accuracy and macro-F1. Removing every proposed component reduces performance, with the largest degradation caused by replacing HPE with a randomly initialized two-layer response-processing transformer. The page-7 early-detection plots additionally show that RPL-Bre outperforms Soft-Ppt, PLAN, STANKER, BiGCN, and RvNN across the tested lifecycle checkpoints.

    TWITTER source WEIBO source
    Model on CatAr-COVID19 Acc. Mac-F1 Acc. Mac-F1
    RPL-Bre 0.631 0.617 0.672 0.664
    RPL-Bre w/o RR 0.605 0.598 0.613 0.611
    RPL-Bre w/o APP 0.622 0.607 0.626 0.624
    RPL-Bre w/o RPP 0.610 0.601 0.633 0.632
    RPL-Bre w/o ViRA 0.626 0.612 0.644 0.634
    RPL-Bre w/o HPE 0.571 0.451 0.581 0.433
    RPL-Bre w/o PV 0.592 0.589 0.621 0.617

    RR denotes response ranking, APP absolute propagation position, RPP relative propagation position, ViRA virtual response augmentation, HPE hierarchical prompt encoding, and PV the prototypical verbalizer. The response-ranking ablation indicates that community responses contribute collective information; both propagation-position ablations show that global depth and local relations are complementary; ViRA improves robustness to noisy or truncated response context; and PV is important for language- and domain-transferable label mapping. In the early-detection experiments, the RPL curve reaches saturated performance after approximately 20 available posts on CatAr-COVID19 and approximately 4 hours on Twitter-COVID19, indicating that the model can make useful predictions before the full propagation thread is observed. A separate layer-depth analysis found the best transfer performance at k=6k=6 frozen SynEncoder layers; using only four layers retained language-specific surface bias, while increasing kk beyond six reduced the number of semantically initialized trainable layers and led to a fluctuating decline.

Coverage note — The paper’s future-work proposals and the full implementation details of every comparison baseline were omitted because they are secondary to the contributed RPL method, its ablations, and its demonstrated zero-shot results.

References

  1. 1.Alam, F.; Shaar, S.; Dalvi, F.; Sajjad, H.; Nikolov, A.; Mubarak, H.; Da San Martino, G.; Abdelali, A.; Durrani, N.; Darwish, K.; et al. 2021. Fighting the COVID-19 Infodemic: Modeling the Perspective of Journalists, Fact-Checkers, Social Media Platforms, Policy Makers, and the Society. In Findings of the Association for Computational Linguistics: EMNLP 2021, 611–649.
  2. 2.Allport, G. W.; and Postman, L. J. 1945. Section of psychology: The basic psychology of rumor. Transactions of the New York Academy of Sciences, 8(2 Series II): 61–81.
  3. 3.Ba, J. L.; Kiros, J. R.; and Hinton, G. E. 2016. Layer normalization. arXiv preprint arXiv:1607.06450.
  4. 4.Bian, T.; Xiao, X.; Xu, T.; Zhao, P.; Huang, W.; Rong, Y.; and Huang, J. 2020. Rumor Detection on Social Media with Bi-Directional Graph Convolutional Networks. In Proceedings of the AAAI Conference on Artificial Intelligence.
  5. 5.Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020. Language models are few-shot learners. Advances in neural information processing systems.
  6. 6.Castillo, C.; Mendoza, M.; and Poblete, B. 2011. Information credibility on twitter. In Proceedings of the 20th international conference on World wide web, 675–684.
  7. 7.Collobert, R.; Weston, J.; Bottou, L.; Karlen, M.; Kavukcuoglu, K.; and Kuksa, P. 2011. Natural language processing (almost) from scratch. Journal of machine learning research, 12(ARTICLE): 2493–2537.
  8. 8.De, A.; Bandyopadhyay, D.; Gain, B.; and Ekbal, A. 2021. A Transformer-Based Approach to Multilingual Fake News Detection in Low-Resource Languages. Transactions on Asian and Low-Resource Language Information Processing.
  9. 9.Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171–4186.
  10. 10.Du, J.; Dou, Y.; Xia, C.; Cui, L.; Ma, J.; and Philip, S. Y. 2021. Cross-lingual COVID-19 Fake News Detection. In 2021 International Conference on Data Mining Workshops (ICDMW), 859–862. IEEE.
  11. 11.Friggeri, A.; Adamic, L.; Eckles, D.; and Cheng, J. 2014. Rumor cascades. In Eighth international AAAI conference on weblogs and social media.
  12. 12.Gehring, J.; Auli, M.; Grangier, D.; Yarats, D.; and Dauphin, Y. N. 2017. Convolutional sequence to sequence learning. In International conference on machine learning. PMLR.
  13. 13.Guo, H.; Cao, J.; Zhang, Y.; Guo, J.; and Li, J. 2018. Rumor detection with hierarchical social attention network. In Proceedings of the 27th ACM International Conference on Information and Knowledge Management, 943–951.
  14. 14.Hannak, A.; Margolin, D.; Keegan, B.; and Weber, I. 2014. Get Back! You Don’t Know Me Like That: The Social Mediation of Fact Checking Interventions in Twitter Conversations. In ICWSM.
  15. 15.He, J.; Zhou, C.; Ma, X.; Berg-Kirkpatrick, T.; and Neubig, G. 2021. Towards a Unified View of Parameter-Efficient Transfer Learning. In International Conference on Learning Representations.
  16. 16.Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; and Gelly, S. 2019. Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning, 2790–2799. PMLR.
  17. 17.Huang, L.; Ma, S.; Zhang, D.; Wei, F.; and Wang, H. 2022. Zero-shot Cross-lingual Transfer of Prompt-based Tuning with a Unified Multilingual Prompt. arXiv preprint arXiv:2202.11451.
  18. 18.Jawahar, G.; Sagot, B.; and Seddah, D. 2019. What does BERT learn about the structure of language? In ACL 2019-57th Annual Meeting of the Association for Computational Linguistics.
  19. 19.Ke, L.; Chen, X.; Lu, Z.; Su, H.; and Wang, H. 2020. A novel approach for cantonese rumor detection based on deep neural network. In 2020 IEEE International Conference on Systems, Man, and Cybernetics (SMC), 1610–1615. IEEE.
  20. 20.Khoo, L. M. S.; Chieu, H. L.; Qian, Z.; and Jiang, J. 2020. Interpretable rumor detection in microblogs by attending to user interactions. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, 8783–8790.
  21. 21.Kwon, S.; Cha, M.; Jung, K.; Chen, W.; and Wang, Y. 2013. Prominent features of rumor propagation in online social media. In 2013 IEEE 13th International Conference on Data Mining, 1103–1108. IEEE.
  22. 22.Lee, C.-H.; Cheng, H.; and Ostendorf, M. 2021. Dialogue State Tracking with a Language Model using Schema-Driven Prompting. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 4937–4949.
  23. 23.Lester, B.; Al-Rfou, R.; and Constant, N. 2021. The Power of Scale for Parameter-Efficient Prompt Tuning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 3045–3059.
  24. 24.Li, X. L.; and Liang, P. 2021. Prefix-Tuning: Optimizing Continuous Prompts for Generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing.
  25. 25.Lin, H.; Ma, J.; Chen, L.; Yang, Z.; Cheng, M.; and Guang, C. 2022. Detect Rumors in Microblog Posts for Low-Resource Domains via Adversarial Contrastive Learning. In Findings of the Association for Computational Linguistics: NAACL 2022.
  26. 26.Lin, H.; Ma, J.; Cheng, M.; Yang, Z.; Chen, L.; and Chen, G. 2021a. Rumor detection on twitter with claim-guided hierarchical graph attention networks. In Proceedings of the 2021 conference on empirical methods in natural language processing (EMNLP).
  27. 27.Lin, H.; Yan, Y.; and Chen, G. 2021. Boosting Low-Resource Intent Detection with in-Scope Prototypical Networks. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).
  28. 28.Lin, X. V.; Mihaylov, T.; Artetxe, M.; Wang, T.; Chen, S.; Simig, D.; Ott, M.; Goyal, N.; Bhosale, S.; Du, J.; et al. 2021b. Few-shot Learning with Multilingual Language Models. arXiv preprint arXiv:2112.10668.
  29. 29.Liu, P.; Yuan, W.; Fu, J.; Jiang, Z.; Hayashi, H.; and Neubig, G. 2021a. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. arXiv preprint arXiv:2107.13586.
  30. 30.Liu, X.; Ji, K.; Fu, Y.; Du, Z.; Yang, Z.; and Tang, J. 2021b. P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks. arXiv preprint arXiv:2110.07602.
  31. 31.Liu, X.; Nourbakhsh, A.; Li, Q.; Fang, R.; and Shah, S. 2015. Real-time rumor debunking on twitter. In Proceedings of the 24th ACM International on Conference on Information and Knowledge Management, 1867–1870.
  32. 32.Loshchilov, I.; and Hutter, F. 2018. Decoupled Weight Decay Regularization. In International Conference on Learning Representations.
  33. 33.Ma, J.; Gao, W.; Mitra, P.; Kwon, S.; Jansen, B. J.; Wong, K.-F.; and Cha, M. 2016. Detecting rumors from microblogs with recurrent neural networks. In Proceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence, 3818–3824.
  34. 34.Ma, J.; Gao, W.; and Wong, K.-F. 2017. Detect Rumors in Microblog Posts Using Propagation Structure via Kernel Learning. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics.
  35. 35.Ma, J.; Gao, W.; and Wong, K.-F. 2018. Rumor Detection on Twitter with Tree-structured Recursive Neural Networks. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics.
  36. 36.Min, S.; Lewis, M.; Hajishirzi, H.; and Zettlemoyer, L. 2022. Noisy Channel Language Model Prompting for Few-Shot Text Classification. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics.
  37. 37.Miyato, T.; Maeda, S.-i.; Koyama, M.; and Ishii, S. 2018. Virtual adversarial training: a regularization method for supervised and semi-supervised learning. IEEE transactions on pattern analysis and machine intelligence.
  38. 38.Rao, D.; Miao, X.; Jiang, Z.; and Li, R. 2021. STANKER: Stacking Network based on Level-grained Attention-masked BERT for Rumor Detection on Social Media. In EMNLP.
  39. 39.Rozsa, A.; Rudd, E. M.; and Boult, T. E. 2016. Adversarial diversity and hard positive generation. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, 25–32.
  40. 40.Schucher, N.; Reddy, S.; and de Vries, H. 2021. The Power of Prompt Tuning for Low-Resource Semantic Parsing. arXiv preprint arXiv:2110.08525.
  41. 41.Schwarz, S.; Theophilo, A.; and Rocha, A. 2020. Emet: Embeddings from multilingual-encoder transformer for fake news detection. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing.
  42. 42.Seoh, R.; Birle, I.; Tak, M.; Chang, H.-S.; Pinette, B.; and Hough, A. 2021. Open Aspect Target Sentiment Classification with Natural Language Prompts. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 6311–6322.
  43. 43.Shaw, P.; Uszkoreit, J.; and Vaswani, A. 2018. Self-Attention with Relative Position Representations. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers).
  44. 44.Snell, J.; Swersky, K.; and Zemel, R. 2017. Prototypical networks for few-shot learning. In Proceedings of the 31st International Conference on Neural Information Processing Systems, 4080–4090.
  45. 45.Tian, L.; Zhang, X.; and Lau, J. H. 2021. Rumour Detection via Zero-Shot Cross-Lingual Transfer Learning. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, 603–618. Springer.
  46. 46.Winata, G. I.; Madotto, A.; Lin, Z.; Liu, R.; Yosinski, J.; and Fung, P. 2021. Language Models are Few-shot Multilingual Learners. In Proceedings of the 1st Workshop on Multilingual Representation Learning, 1–15.
  47. 47.Wu, K.; Yang, S.; and Zhu, K. Q. 2015. False rumors detection on sina weibo by propagation structures. In 2015 IEEE 31st international conference on data engineering. IEEE.
  48. 48.Yang, F.; Liu, Y.; Yu, X.; and Yang, M. 2012. Automatic detection of rumor on sina weibo. In Proceedings of the ACM SIGKDD workshop on mining data semantics, 1–7.
  49. 49.Yang, Z.; Ma, J.; Chen, H.; Lin, H.; Luo, Z.; and Chang, Y. 2022. A Coarse-to-fine Cascaded Evidence-Distillation Neural Network for Explainable Fake News Detection. In Proceedings of the 29th International Conference on Computational Linguistics, 2608–2621.
  50. 50.Yao, Y.; Rosasco, L.; and Caponnetto, A. 2007. On early stopping in gradient descent learning. Constructive Approximation, 26(2): 289–315.
  51. 51.Yu, F.; Liu, Q.; Wu, S.; Wang, L.; and Tan, T. 2017. A Convolutional Approach for Misinformation Identification. In IJCAI.
  52. 52.Zhao, M.; and Schutze, H. 2021. Discrete and Soft Prompting for Multilingual Models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing.
  53. 53.Zhao, Z.; Resnick, P.; and Mei, Q. 2015. Enquiring minds: Early detection of rumors in social media from enquiry posts. In Proceedings of the 24th international conference on world wide web, 1395–1405.
  54. 54.Zubiaga, A.; Aker, A.; Bontcheva, K.; Liakata, M.; and Procter, R. 2018. Detection and resolution of rumours in social media: A survey. ACM Computing Surveys (CSUR).

Citation

MLA
Lin, H., et al. “Zero-Shot Rumor Detection with Propagation Structure via Prompt Learning”. arXiv, 2022, http://arxiv.org/abs/2212.01117v5.
APA
Lin, H., Yi, P., Ma, J., Jiang, H., Luo, Z., Shi, S., & Liu, R. (2022). Zero-Shot Rumor Detection with Propagation Structure via Prompt Learning. arXiv. http://arxiv.org/abs/2212.01117v5
Chicago
Lin, H., P. Yi, J. Ma, et al. 2022. “Zero-Shot Rumor Detection with Propagation Structure via Prompt Learning”. arXiv. http://arxiv.org/abs/2212.01117v5.
Harvard
Lin, H. et al. (2022) “Zero-Shot Rumor Detection with Propagation Structure via Prompt Learning”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2212.01117v5.
Vancouver
1. Lin H, Yi P, Ma J, Jiang H, Luo Z, Shi S, Liu R (2022) Zero-Shot Rumor Detection with Propagation Structure via Prompt Learning. arXiv

BibTeX

@article{lin2022zero,
  title = {Zero-Shot Rumor Detection with Propagation Structure via Prompt Learning},
  author = {Lin, Hongzhan and Yi, Pengyao and Ma, Jing and Jiang, Haiyun and Luo, Ziyang and Shi, Shuming and Liu, Ruifang},
  year = {2022},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2212.01117v5},
  eprint = {2212.01117}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF