Information credibility on twitter

Carlos CastilloMarcelo MendozaBárbara Poblete

article2011WWW2,620 citationsSeoul Test of Time Award

Demonstrates that machine learning classifiers can automatically evaluate the credibility of breaking news on Twitter by extracting measurable signals from tweet text, user behavior, external citations, and propagation trees.

Listen

Twitter spreads both accurate news and false rumors rapidly, especially during crises when official information is scarce. Users struggle to judge credibility without traditional cues, and misinformation can heighten public anxiety.

The article set out to test whether signals available inside the Twitter platform itself can support automatic classification of trending topics as newsworthy and, separately, as credible or not credible.

Researchers collected more than 2,500 bursty topics over two months using an online trend detector. They obtained human labels through Mechanical Turk for 383 topics on whether each was news or chat, then applied a supervised classifier to identify 747 newsworthy cases. A second round of Mechanical Turk assessments produced credibility labels for those cases. The team extracted dozens of features covering message content, user profiles, topic aggregates, and retweet propagation patterns, then trained decision-tree classifiers.

The newsworthiness classifier reached 89 percent accuracy. The credibility classifier, distinguishing almost-certainly-true items from the rest, reached 86 percent accuracy. Credible topics tended to contain URLs, originate from active users with many followers, show deeper propagation trees, and carry more negative sentiment; non-credible topics more often featured questions, positive sentiment, emoticons, and single-level bursts.

These measurable differences indicate that platform-internal signals can approximate human credibility judgments at useful levels of precision. Automatic filters could therefore help users, journalists, and emergency responders focus on higher-quality information and reduce exposure to rumors.

The authors recommend extending the approach to larger and partial data streams, examining the content of linked pages, and testing additional context signals such as poster follower counts and avatars. Further validation on non-trending topics and different languages is needed before operational deployment.

The study relies on Mechanical Turk labels that may contain noise, covers only English-language trending topics from a two-month window, and excludes ambiguous cases; results should therefore be treated as promising but preliminary indicators rather than definitive performance guarantees.

  • Paper: Fake News Detection on Social Media: A Data Mining Perspective, Kai Shu et al. (2017). Surveys the broader landscape of automated fake news and misinformation detection on social media, synthesizing the feature-based and social-context approaches introduced by early credibility studies.
  • Paper: The rise of social bots, Emilio Ferrara et al. (2014). Investigates the behavior and detection of automated social bots that actively manipulate credibility, engagement, and information diffusion across social platforms.
  • Paper: FEVER: a Large-scale Dataset for Fact Extraction and VERification, James Thorne et al. (2018). Advances automated veracity assessment by constructing a large-scale benchmark that couples claim classification directly with evidence retrieval and verification.
Cover for Information credibility on twitter

Abstract

We analyze the information credibility of news propagated through Twitter, a popular microblogging service. Previous research has shown that most of the messages posted on Twitter are truthful, but the service is also used to spread misinformation and false rumors, often unintentionally.

On this paper we focus on automatic methods for assessing the credibility of a given set of tweets. Specifically, we analyze microblog postings related totrendingtopics, and classify them as credible or not credible, based on features extracted from them. We use features from usersposting and re-posting (“re-tweeting”) behavior, from the text of the posts, and from citations to external sources.

We evaluate our methods using a significant number of human assessments about the credibility of items on a recent sample of Twitter postings. Our results shows that there are measurable differences in the way messages propagate, that can be used to classify them automatically as credible or not credible, with precision and recall in the range of 70% to 80%.

Table of Contents

  • 1. INTRODUCTION
  • 2. RELATED WORK
  • 3. DATA COLLECTION
  • 3.1 Automatic event detection
  • 3.2 Newsworthy topic assessments
  • 3.3 Credibility assessment
  • 4. AUTOMATICCREDIBILITYANALYSIS
  • 4.1 Social media credibility
  • 4.2 Automatically finding newsworthy topics
  • 4.3 Feature analysis for the credibility task
  • 4.4 Automatically assessing credibility
  • 4.5 Credibility analysis at feature-level
  • 5. CONCLUSIONS
  • 6. REFERENCES

Knowls

  1. Knowl 1 — Supervised Credibility Prediction on Twitter Topics

    empirical result

    A supervised J48 decision tree classifier trained on newsworthy Twitter topics successfully predicts credibility levels, discriminating between credible news (Class A: "almost certainly true") and non-credible rumors (Class B: "likely to be false" and "almost certainly false").

    The evaluation was conducted on a balanced set of 608 labeled news topics (306 Class A instances and 302 Class B instances; 50.3% / 49.7% balance) using 3-fold cross-validation.

    Overall performance metrics:

    • Correctly Classified Instances: 86.0119%
    • Kappa Statistic: 0.7189
    • Mean Absolute Error (MAE): 0.1540
    • Root Mean Squared Error (RMSE): 0.3608
    • Relative Absolute Error: 30.8711%
    • Root Relative Squared Error: 72.2466%
    Class TP Rate FP Rate Precision Recall F1F_1
    A ("true") 0.825 0.108 0.874 0.825 0.849
    B ("false") 0.892 0.175 0.849 0.892 0.870
    Weighted Avg. 0.860 0.143 0.861 0.860 0.860

    The classifier achieves balanced precision and recall across both credible and non-credible classes, substantially outperforming random baseline guessing (as shown by a Kappa statistic of 0.7189).

  2. Knowl 2 — Predictive Signals and Decision Rules for Credibility on Microblogs

    empirical result

    Analysis of the GINI split criteria in the J48 decision tree classifier reveals clear structural and behavioral rules that separate credible news events from false rumors on microblogging platforms:

    1. URL Inclusion: The fraction of tweets containing a URL is the root decision node of the tree. Topics where tweets consistently lack URLs strongly predict non-credible news.
    2. Sentiment Polarity: Topics with a higher fraction of negative sentiment words correlate with credible news events. Conversely, topics characterized by a higher fraction of positive sentiment words, smiling emoticons, or question marks are associated with non-credible information and rumors.
    3. User History and Activity: Credible news is propagated predominantly by users who have higher historical status counts (more past tweets authored) and established social networks (higher follower and friend counts), whereas low-credibility news is propagated by less active users.
    4. Propagation Tree Structure: Credible news exhibits significantly larger maximum level sizes in its retweet propagation trees, indicating wide multi-level rebroadcasting and community propagation.
  3. Knowl 3 — Feature Hierarchy for Social Media Credibility Assessment

    model/method

    A four-tier feature engineering hierarchy characterizes microblog discussion topics across different analytical scopes for credibility assessment:

    1. Message-based Features: Properties extracted directly from individual tweet texts, including character length, word count, presence of question marks, exclamation marks, emoticons (smile/frown), 1st/2nd/3rd person pronouns, uppercase character fraction, URL counts, citations to top 100/1,000/10,000 domains, user mentions, hashtags, stock symbols, retweet indicators, posting day of the week, positive/negative sentiment words, and overall sentiment score.
    2. User-based Features: Metadata reflecting author reputation and platform tenure, including account registration age (in days), statuses count (historical tweet volume), follower count, friend (followee) count, verified account flag, presence of a profile biography, and presence of an external URL in the user profile.
    3. Topic-based Features: Aggregated statistical metrics computed across all tweets belonging to a topic cluster, such as the fraction of tweets containing URLs, hashtags, user mentions, question marks, exclamation marks, emoticons, pronouns, or positive/negative sentiment; average tweet length; average author registration age; average status count; average follower/friend counts; fraction of verified authors; and concentration metrics measuring the fraction of topic volume contributed by the single most frequent URL, hashtag, user mention, or author.
    4. Propagation-based Features: Topological properties of the retweet trees formed around the topic, including the degree of the root (number of initial tweets), maximum subtree size, maximum and average node degrees (excluding root), maximum and average node depth, and maximum level size (maximum node count at any depth level excluding direct children of the root).
  4. Knowl 4 — Feature Subset Performance for Credibility Classification

    empirical result

    Evaluating four distinct subsets of features using 3-fold cross-validated J48 decision trees demonstrates the varying predictive utility of different information channels for social media credibility classification:

    Feature Subset Class TP Rate FP Rate Precision Recall F1F_1
    Text (20 features) A 0.636 0.152 0.808 0.636 0.712
    B 0.848 0.364 0.700 0.848 0.767
    W. Avg. 0.742 0.258 0.754 0.742 0.739
    Network (7 features) A 0.667 0.212 0.759 0.667 0.710
    B 0.788 0.333 0.703 0.788 0.743
    W. Avg. 0.727 0.273 0.731 0.727 0.726
    Propagation (6 features) A 0.606 0.091 0.870 0.606 0.714
    B 0.909 0.394 0.698 0.909 0.789
    W. Avg. 0.758 0.242 0.784 0.758 0.752
    Top-element (4 features) A 0.727 0.152 0.828 0.727 0.774
    B 0.848 0.273 0.757 0.848 0.800
    W. Avg. 0.788 0.212 0.792 0.788 0.787

    Key takeaways:

    • The Propagation subset achieves the highest true positive rate for detecting false news (Class B TP Rate = 0.909), showing that diffusion topology is critical for identifying rumors.
    • The Top-element subset (shares of the top URL, hashtag, mention, and author) achieves the highest overall weighted F1F_1 score (0.787) and the highest F1F_1 for credible news (0.774), demonstrating that source concentration patterns are highly informative.
    • Text-only and network-only feature sets underperform, indicating that textual content and author network metrics alone are insufficient for credibility assessment without propagation and concentration signals.
  5. Knowl 5 — Supervised Classification of Newsworthy Topics from Microblog Streams

    empirical result

    A supervised cost-sensitive learning approach automatically distinguishes bursty discussion topics on microblogs as newsworthy events (NEWS), informal conversations (CHAT), or uncertain (UNSURE).

    The classifier employs a J48 decision tree trained with a cost matrix assigning misclassification costs involving the NEWS class as 1.01.0, and misclassifications between CHAT and UNSURE as 0.50.5. Training uses a 300% bootstrapping strategy with replacement under a uniform prior across classes, evaluated under 3-fold cross-validation.

    Overall performance metrics:

    • Accuracy: 89.121%
    • Kappa Statistic: 0.8368
    • Mean Absolute Error (MAE): 0.0806
    • Root Mean Squared Error (RMSE): 0.2569
    • Relative Absolute Error: 18.1388%
    • Root Relative Squared Error: 54.4912%
    Class TP Rate FP Rate Precision Recall F1F_1
    NEWS 0.927 0.039 0.922 0.927 0.924
    CHAT 0.874 0.054 0.892 0.874 0.883
    UNSURE 0.873 0.070 0.860 0.873 0.866
    Weighted Avg. 0.891 0.054 0.891 0.891 0.891

    The model achieves a true positive rate of 0.927 and an F1F_1 score of 0.924 for identifying newsworthy topics while maintaining a false positive rate of 0.039.

  6. Knowl 6 — Two-Stage Microblog Information Credibility Framework

    model/method

    The automated framework for evaluating information credibility in social media streams operates via a two-stage pipeline:

    1. Topic and Event Extraction: Burst detection over the microblog stream identifies keyword clusters (AB)(A \wedge B) where AA is a conjunction of keywords/hashtags and BB is a disjunction. All messages matching the query in a two-day window centered on the burst peak constitute a candidate topic.
    2. Stage 1 — Newsworthiness Filtering: A cost-sensitive supervised classifier filters candidate topics, categorizing them into NEWS (broad public relevance) versus CHAT / UNSURE (personal chatter and informal opinions).
    3. Stage 2 — Credibility Assessment: For all topics classified as NEWS, features across message, user, topic, and propagation dimensions are extracted and evaluated by a credibility classifier to predict whether the topic is credible ("almost certainly true") or non-credible ("likely/almost certainly false").
  7. Knowl 7 — Definition of Social Media Credibility

    definition

    Social media credibility is defined as the dimension of information credibility—understood in the sense of believability, or offering reasonable grounds for being believed—that can be assessed exclusively from internal signals, metadata, behavioral interactions, textual cues, and propagation dynamics observable within the social media platform itself, without reliance on external ground-truth verification.

  8. Knowl 8 — Crowdsourced Dataset Annotation Protocol for Microblog Credibility

    experimental setup

    Ground-truth credibility labels for microblog topics are established using a two-phase crowdsourced annotation methodology via Amazon Mechanical Turk:

    1. Newsworthiness Labeling: Evaluators are presented with 10 sample tweets and burst keywords for a topic. Each task is assessed by 7 independent annotators who must provide a short descriptive summary sentence (used to filter click-spam). A label of NEWS or CHAT is assigned if at least 5 out of 7 annotators agree; otherwise, it is labeled UNSURE. Out of 383 randomly sampled burst topics, 29.5% (113) were labeled NEWS, 34.9% (134) CHAT, and 35.6% (136) UNSURE.
    2. Credibility Labeling: For topics classified as NEWS, annotators review 10 sample tweets and a descriptive topic sentence, categorizing credibility as "almost certainly true", "likely to be false", "almost certainly false", or "I can't decide" (with required justification sentences). A class label requires agreement of at least 5 out of 7 judges. Across 747 assessed news topics, 41.0% (306) were "almost certainly true" (Class A), 31.8% (237) "likely to be false", 8.6% (65) "almost certainly false" (yielding 302 Class B cases), and 18.6% (139) ambiguous/undecided.

Coverage note — No substantial contributed material was omitted from the knowls.

References

  1. 1.E. Agichtein, C. Castillo, D. Donato, A. Gionis, and G. Mishne. Finding high-quality content in social media. In WSDM '08: Proceedings of the international conference on Web search and web data mining, pages 183–194, New York, NY, USA, 2008. ACM.
  2. 2.Alonso, Omar, Carson, Chad, Gerster, David, Ji, Xiang, and Nabar, Shubha. Detecting Uninteresting Content in Text Streams. In SIGIR Crowdsourcing for Search Evaluation Workshop, 2010.
  3. 3.C. L. Armstrong and M. J. Mcadams. Blogs of information: How gender cues and individual motivations influence perceptions of credibility. Journal of Computer-Mediated Communication, 14(3):435–456, 2009.
  4. 4.F. Benevenuto, G. Magno, T. Rodrigues, and V. Almeida. Detecting Spammers on Twitter. In Collaboration, Electronic messaging, Anti-Abuse and Spam Conference (CEAS), July 2010.
  5. 5.R. Crane and D. Sornette. Robust dynamic classes revealed by measuring the response function of a social system. Proceedings of the National Academy of Sciences, 105(41):15649–15653, October 2008.
  6. 6.B. De Longueville, R. S. Smith, and G. Luraschi. "OMG, from here, I can see the flames!": a use case of mining location based social networks to acquire spatio-temporal data on forest fires. In LBSN '09: Proceedings of the 2009 International Workshop on Location Based Social Networks, pages 73–80, New York, NY, USA, 2009. ACM.
  7. 7.P. S. Earle, M. Guy, C. Ostrum, S. Horvath, and R. A. Buckmaster. OMG Earthquake! Can Twitter improve earthquake response? AGU Fall Meeting Abstracts, pages B1697+, Dec. 2009.
  8. 8.A. J. Flanagin and M. J. Metzger. Perceptions of internet information credibility. Journalism and Mass Communication Quarterly, 77(3):515–540, 2000.
  9. 9.A. J. Flanagin and M. J. Metzger. The role of site features, user attributes, and information verification behaviors on the perceived credibility of web-based information. New Media Society, 9(2):319–342, April 2007.
  10. 10.B. J. Fogg and H. Tseng. The elements of computer credibility. In CHI '99: Proceedings of the SIGCHI conference on Human factors in computing systems, pages 80–87, New York, NY, USA, 1999. ACM.
  11. 11.C. Grier, K. Thomas, V. Paxson, and M. Zhang. @spam: the underground on 140 characters or less. In CCS '10: Proceedings of the 17th ACM conference on Computer and Communications Security, CCS '10, pages 27–37, New York, NY, USA, October 2010. ACM.
  12. 12.A. L. Hughes and L. Palen. Twitter adoption and use in mass convergence and emergency events. In ISCRAM Conference, May 2009.
  13. 13.A. Java, X. Song, T. Finin, and B. Tseng. Why we twitter: understanding microblogging usage and communities. In WebKDD/SNA-KDD '07: Proceedings of the 9th WebKDD and 1st SNA-KDD 2007 workshop on Web mining and social network analysis, pages 56–65, New York, NY, USA, 2007. ACM.
  14. 14.T. J. Johnson, B. K. Kaye, S. L. Bichard, and W. J. Wong. Every blog has its day: Politically-interested internet users' perceptions of blog credibility. Journal of Computer-Mediated Communication, 13(1), 2007.
  15. 15.K. Kireyev, L. Palen, and K. Anderson. Applications of topics models to analysis of disaster-related twitter data. In NIPS Workshop on Applications for Topic Models: Text and Beyond, December 2009.
  16. 16.H. Kwak, C. Lee, H. Park, and S. Moon. What is twitter, a social network or a news media? In World Wide Web Conference. ACM Press, 2010.
  17. 17.V. Lampos, T. D. Bie, and N. Cristianini. Flu detector - tracking epidemics on twitter. In European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD 2010), pages 599–602, Barcelona, Spain, 2010. Springer, Springer.
  18. 18.M. Mathioudakis and N. Koudas. TwitterMonitor: trend detection over the twitter stream. In Proceedings of the 2010 international conference on Management of data, pages 1155–1158. ACM, 2010.
  19. 19.M. Mendoza, B. Poblete, and C. Castillo. Twitter under crisis: Can we trust what we rt? In 1st Workshop on Social Media Analytics (SOMA '10). ACM Press, July 2010.
  20. 20.E. Mustafaraj and P. Metaxas. From obscurity to prominence in minutes: Political speech and real-time search. In Proceedings of the WebSci10: Extending the Frontiers of Society On-Line, April 2010.
  21. 21.M. Naaman, J. Boase, and C. H. Lai. Is it really about me?: message content in social awareness streams. In Proceedings of the 2010 ACM conference on Computer supported cooperative work, CSCW '10, pages 189–192, New York, NY, USA, 2010. ACM.
  22. 22.Pear Analytics. Twitter study. http://www.pearanalytics.com/wp-content/uploads/2009/-08/Twitter-Study-August-2009.pdf, August 2009.
  23. 23.Pew Research Center. Internet Overtakes Newspapers As News Outlet. http://pewresearch.org/pubs/1066/internet-overtakes-newspapers-as-news-source 2008.
  24. 24.A. M. Popescu and M. Pennacchiotti. Detecting controversial events from twitter. In Proceedings of the 19th ACM international conference on Information and knowledge management, CIKM '10, pages 1873–1876, New York, NY, USA, 2010. ACM.
  25. 25.K. Poulsen. Firsthand reports from california wildfires pour through twitter. October 2007.
  26. 26.J. Ratkiewicz, M. Conover, M. Meiss, B. Gonçalves, S. Patil, A. Flammini, and F. Menczer. Detecting and Tracking the Spread of Astroturf Memes in Microblog Streams. arXiv, Nov 2010.
  27. 27.T. Sakaki, M. Okazaki, and Y. Matsuo. Earthquake shakes Twitter users: real-time event detection by social sensors. In Proceedings of the 19th international conference on World wide web, WWW '10, pages 851–860, New York, NY, USA, April 2010. ACM.
  28. 28.J. Sankaranarayanan, H. Samet, B. E. Teitler, M. D. Lieberman, and J. Sperling. TwitterStand: news in tweets. In GIS '09: Proceedings of the 17th ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems, pages 42–51, New York, NY, USA, November 2009. ACM Press.
  29. 29.M. Schmierbach and A. Oeldorf-Hirsch. A little bird told me, so i didn't believe it: Twitter, credibility, and issue perceptions. In Proc. of annual meeting of the Association for Education in Journalism and Mass Communication. AEJMC, August 2010.
  30. 30.J. Schwarz and M. R. Morris. Augmenting Web Pages and Search Results to Support Credibility Assessment. In ACM Conference on Human Factors in Computing Systems (CHI). ACM Press, May 2011.
  31. 31.K. Starbird, L. Palen, A. L. Hughes, and S. Vieweg. Chatter on the red: what hazards threat reveals about the social life of microblogged information. In CSCW '10: Proceedings of the 2010 ACM conference on Computer supported cooperative work, pages 241–250, New York, NY, USA, 2010. ACM.
  32. 32.S. Vieweg. Microblogged contributions to the emergency arena: Discovery, interpretation and implications. In Computer Supported Collaborative Work, February 2010.
  33. 33.S. Vieweg, A. Hughes, K. Starbird, and L. Palen. Microblogging during two natural hazards events: What twitter may contribute to situational awareness. In Proceedings of ACM Conference on Computer Human Interaction (CHI), April 2010.
  34. 34.C. R. W. Watch. Leap of faith: Using the internet despite the dangers. http://www.consumerwebwatch.org/pdfs/princeton.pdf, October 2005.
  35. 35.D. J. Watts and J. Peretti. Viral Marketing for the Real World. Harvard Business Review, June 2007.
  36. 36.S. Yardi, D. Romero, G. Schoenebeck, and D. Boyd. Detecting spam in a Twitter network. First Monday, 15(1), January 2010.

Citation

MLA
Castillo, C., et al. “Information Credibility on Twitter”. Proceedings of the 20th International Conference on World Wide Web, 2011, pp. 675–84, https://doi.org/10.1145/1963405.1963500.
APA
Castillo, C., Mendoza, M., & Poblete, B. (2011). Information credibility on twitter. Proceedings of the 20th International Conference on World Wide Web, 675–684. https://doi.org/10.1145/1963405.1963500
Chicago
Castillo, C., M. Mendoza, and B. Poblete. 2011. “Information Credibility on Twitter”. Proceedings of the 20th International Conference on World Wide Web, 675–84. https://doi.org/10.1145/1963405.1963500.
Harvard
Castillo, C., Mendoza, M. and Poblete, B. (2011) “Information credibility on twitter”, Proceedings of the 20th international conference on World wide web. ACM, pp. 675–684. Available at: https://doi.org/10.1145/1963405.1963500.
Vancouver
1. Castillo C, Mendoza M, Poblete B (2011) Information credibility on twitter. In: Proceedings of the 20th international conference on World wide web. ACM, pp 675–684

BibTeX

@inproceedings{Castillo_2011, series={WWW ’11}, title={Information credibility on twitter}, url={http://dx.doi.org/10.1145/1963405.1963500}, DOI={10.1145/1963405.1963500}, booktitle={Proceedings of the 20th international conference on World wide web}, publisher={ACM}, author={Castillo, Carlos and Mendoza, Marcelo and Poblete, Barbara}, year={2011}, month=Mar, pages={675–684}, collection={WWW ’11} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF