Detecting Rumors from Microblogs with Recurrent Neural Networks
Jing MaWei GaoPrasenjit MitraSejeong KwonBernard J. JansenKam-Fai WongMeeyoung Cha
Presents a recurrent neural network framework that models social media posts as sequential time series to learn temporal and contextual representations, enabling faster and more accurate early rumor detection than models reliant on manual feature engineering.
The rapid spread of false rumors across microblogging platforms presents serious real-world risks, often triggering public panic, severe social disruption, and costly emergencies. Conventional detection methods rely heavily on human fact-checking or machine-learning algorithms built on manually engineered features and simple keyword patterns. However, manual verification is slow and limited in scale, while hand-crafted features require painstaking effort, introduce bias, and struggle to capture how public skepticism and evidence evolve over time.
The article demonstrates an automated, deep-learning approach that identifies rumors at the aggregate event level by modeling microblog streams as variable-length time series. The primary objective is to evaluate whether recurrent neural networks can autonomously learn the complex, time-dependent linguistic and contextual signals that distinguish rumors from factual events without relying on manual feature engineering.
To evaluate this approach, the authors constructed and analyzed two large-scale datasets from Twitter and Sina Weibo comprising over 5,000 verified claims and nearly 5 million posts. Microblog posts for each event were segmented into dynamic, equal-duration time intervals based on continuous activity. The system then converted the most relevant vocabulary terms within each interval into continuous representations and processed them through various sequential neural architectures, including basic recurrent units and advanced gated structures with single and multiple layers.
The experimental findings show that deep learning significantly outperforms existing detection methods across all evaluation metrics. Advanced gated architectures achieved the highest overall accuracy, reaching 88.1% on Twitter and 91.0% on Sina Weibo, clearly exceeding the best baseline models (80.8% and 85.7%, respectively). Incorporating gated units and adding a second hidden layer proved especially effective at filtering noise and capturing long-distance dependencies in text streams. Most importantly, the proposed model excelled at early detection: within just 12 hours of initial propagation, it achieved 83.9% accuracy on Twitter and 89.0% on Weibo, identifying rumors far faster and more accurately than existing algorithms and official debunking services.
These results show that automated sequential modeling can reliably detect emerging misinformation directly from evolving public discourse, eliminating the overhead of manual feature design. For social media platforms, safety agencies, and decision-makers, this capability substantially reduces operational response times, mitigating the risks of public panic and misinformation campaigns before they escalate. While the empirical evidence provides strong confidence in the model's accuracy on the evaluated platforms, future work should explore unsupervised learning methods to harness massive unlabeled social streams and further evaluate generalizability across different communication channels.
- Paper: Information credibility on twitter, Carlos Castillo et al. (2011). Establishes the foundational benchmark and hand-crafted feature framework for microblog credibility assessment that the RNN-based rumor detection model is designed to surpass.
- Paper: Long Short-Term Memory, Sepp Hochreiter et al. (1997). Introduces the Long Short-Term Memory architecture essential for modeling the long-distance temporal dependencies and continuous representations of microblog posts.
- Paper: Recurrent Convolutional Neural Networks for Text Classification, Siwei Lai et al. (2015). Demonstrates how recurrent neural architectures capture sequential context for text classification without relying on hand-engineered features.
- Paper: Fake News Detection on Social Media: A Data Mining Perspective, Kai Shu et al. (2017). Synthesizes social media misinformation research into a comprehensive survey, contextualizing deep learning and temporal rumor detection approaches.
- Paper: “Liar, Liar Pants on Fire”: A New Benchmark Dataset for Fake News Detection, William Yang Wang (2017). Builds upon neural text classification for veracity assessment by introducing a large-scale benchmark dataset incorporating metadata.
- Paper: The COVID-19 social media infodemic, Matteo Cinelli et al. (2020). Applies large-scale tracking of rumor and misinformation dynamics across social platforms in a real-world infodemic context.
