The rise of social bots

Emilio FerraraOnur VarolClayton DavisFilippo MenczerAlessandro Flammini

article2014Communications of the ACM2,101 citations

Examines the societal risks posed by deceptive social bots and details the behavioral, network, and temporal signatures used to detect automated manipulation across online platforms.

arXiv: 1407.5225
  • Paper: Fake News Detection on Social Media: A Data Mining Perspective, Kai Shu et al. (2017). Reading this comprehensive survey on fake news detection first establishes the foundational data-mining and social-context frameworks that the source paper expands upon for bot-driven manipulation.
  • Paper: What is Twitter, a social network or a news media?, Haewoon Kwak et al. (2010). Reviewing this early Twitter network analysis provides essential context on platform topology and information diffusion dynamics before examining how automated social bots exploit those pathways.
Cover for The rise of social bots

Abstract

The Turing test aimed to recognize the behavior of a human from that of a computer algorithm. Such challenge is more relevant than ever in today's social media context, where limited attention and technology constrain the expressive power of humans, while incentives abound to develop software agents mimicking humans. These social bots interact, often unnoticed, with real people in social media ecosystems, but their abundance is uncertain. While many bots are benign, one can design harmful bots with the goals of persuading, smearing, or deceiving. Here we discuss the characteristics of modern, sophisticated social bots, and how their presence can endanger online ecosystems and our society. We then review current efforts to detect social bots on Twitter. Features related to content, network, sentiment, and temporal patterns of activity are imitated by bots but at the same time can help discriminate synthetic behaviors from human ones, yielding signatures of engineered social tampering.

Table of Contents

  • References

Knowls

  1. Knowl 1 — Taxonomy of Social Bot Detection Systems

    model/method

    Social bot detection approaches can be classified into three primary paradigms alongside hybrid frameworks:

    1. Graph-based detection: Analyzes the structure of the social graph (e.g., community detection, graph traversal, and tie density) under the hypothesis that bots and Sybils form tightly-knit clusters with sparse connections to legitimate accounts.
    2. Crowd-sourcing systems: Leverages human intelligence through Social Turing Tests, where crowd workers or expert annotators evaluate user profiles and conversational nuances (e.g., sarcasm or linguistic abnormalities) using majority voting schemes.
    3. Feature-based machine learning: Extracts multi-modal behavioral, metadata, linguistic, and temporal features from account activities and trains supervised or semi-supervised classifiers to separate automated accounts from human users.
    4. Hybrid and synchronization-based systems: Integrates network topology, temporal correlation, and clickstream sequences to identify coordinated, synchronous campaign activities across multiple accounts.
  2. Knowl 2 — Feature Classes for Behavioral Social Bot Detection

    model/method

    Feature-based bot detection systems categorize user activity into orthogonal behavioral and structural dimensions:

    • Network features: Derived from interaction graphs (retweets, mentions, and hashtag co-occurrences), measuring topological metrics such as degree distribution, clustering coefficient, and centrality.
    • User metadata features: Account-level attributes provided by the platform, including interface language, geographic locations, profile description, and account creation timestamp.
    • Friend features: Descriptive statistics and distribution moments (mean, median, variance, entropy) calculated over the numbers of followers, followees, and posts of an account's social neighborhood.
    • Timing features: Temporal dynamics of content generation (posting) and consumption (sharing/retweeting), including inter-post time intervals and goodness-of-fit against Poisson processes.
    • Content features: Natural language processing cues derived from message text, such as part-of-speech tagging frequencies (verb, noun, and adverb ratios) and vocabulary distributions.
    • Sentiment features: Emotional valence, arousal, dominance, happiness indices, and emoticon frequencies evaluated across a user's stream of messages.
  3. Knowl 3 — Discriminative Behavioral and Metadata Signatures of Social Bots versus Humans

    empirical result

    Empirical comparison of normalized behavioral metrics (zz-scores) between social bots and human users on microblogging platforms demonstrates clear distinguishing patterns:

    • Retweeting volume: Bots exhibit substantially higher rates of retweeting/rebroadcasting content compared to humans.
    • Username length: Bot accounts have systematically longer usernames on average than legitimate human accounts.
    • Account age: Bot accounts tend to be significantly younger/more recently created than human accounts.
    • Original content and interaction: Bots generate fewer original tweets, direct replies, and user mentions than humans.
    • Inbound influence: Bots are retweeted and cited by other users significantly less frequently than legitimate human accounts.
  4. Knowl 4 — Definition of Social Bot

    definition

    A social bot is an automated computer algorithm that operates on social media platforms to generate content and interact with human users. Its objective is typically to emulate human behavior, infiltrate online networks, and potentially influence public discourse, spread misinformation, alter sentiment, or manipulate platform analytics across dimensions including message content, temporal posting patterns, social connectivity, and information diffusion.

  5. Knowl 5 — Feature-Based Detection Framework and Bot or Not? Architecture

    model/method

    The Bot or Not? architecture is a supervised machine learning detection system for social bots on Twitter. It operates by extracting over a thousand predictive features spanning network topology, user metadata, friend statistics, timing dynamics, linguistic content, and sentiment scores. When trained on benchmark datasets (such as the Texas A&M honeypot dataset comprising approximately 15,000 bot accounts and 15,000 human accounts along with millions of tweets), standard supervised classifiers achieve an Area Under the Receiver Operating Characteristic curve (AUROC\mathrm{AUROC}) exceeding 0.950.95 via cross-validation.

  6. Knowl 6 — Failure Modes of Innocent-by-Association Graph-Based Sybil Defense

    limitation

    Graph-based detection frameworks (e.g., SybilRank) frequently rely on the innocent-by-association assumption, positing that accounts establishing links with legitimate users are themselves legitimate, and that legitimate users reject connections from unknown accounts. This assumption fails in practice because:

    • Empirical measurements show that over 20% of social media users accept friendship requests indiscriminately, and over 60% accept requests from accounts sharing at least one mutual contact.
    • Microblogging platforms (e.g., Twitter, Tumblr) are fundamentally structured around following and interacting with strangers.
    • Sophisticated bots can forge realistic community structures and infiltrate legitimate clusters, leading innocent-by-association methods to produce high false positive rates, which risk erroneously suspending genuine human users.
  7. Knowl 7 — Synchronized Behavioral and Clickstream Analysis for Sybil Detection

    model/method

    Detection frameworks combining behavioral clickstream analysis and coordinated activity recognition (e.g., Renren Sybil detector, CopyCatch, SynchroTrap) detect botnets by capturing synchronized actions across user populations:

    • Clickstream differentiation: Legitimate users spend comparatively more session time browsing content (such as photos and videos) and messaging friends, whereas Sybil accounts spend session time harvesting profiles, sending friend requests, and colluding.
    • Windowed event profiling: Highly accurate classification can be achieved using only a sequence of the last 100 click events per user, avoiding the overhead of storing full historical clickstreams.
    • "Sybil until proven otherwise" paradigm: Rather than presuming innocence by association, these models identify dense temporal and content-similarity clusters among colluding accounts, allowing detection of novel evasion strategies (such as text embedded within images) while maintaining low false-positive rates.
  8. Knowl 8 — Operational Constraints of Crowd-Sourced Social Bot Detection

    limitation

    Deploying human crowd-sourcing and Social Turing Tests for social bot detection exhibits three major operational bottlenecks:

    1. Performance degradation: Individual detection accuracy among hired crowd workers deteriorates over time, necessitating expert analysts to maintain ground-truth quality.
    2. Economic scalability: Employing human annotators and majority voting protocols across platforms with hundreds of millions of users is economically infeasible compared to automated algorithmic pipelines.
    3. Privacy risks: Exposing user profiles, direct interactions, and behavioral data to third-party distributed workers introduces significant privacy and compliance concerns.
  9. Knowl 9 — Adversarial Arms Race and Detection Limitations for Hybrid Accounts

    limitation

    Supervised social bot detection faces inherent vulnerabilities against evolving adversarial strategies:

    • Cyborgs and compromised accounts: Feature-based classifiers fail when evaluating cyborg accounts (which blend automated scheduling with manual human control) or hijacked legitimate accounts, as their behavioral signatures match human baselines.
    • Smoke-screening tactics: Modern botnets frequently abandon individual human mimicry in favor of synchronized smoke-screening—flooding conversational channels with noise and coordination to distract from specific topics.
    • Non-stationarity and active learning: Static supervised models degrade as bot creators adopt realistic circadian rhythms and advanced natural language generation, requiring continuous retraining and active learning paradigms to elevate the cost of deception.

Coverage note — All primary survey contributions—including the detection taxonomy, feature engineering classes, empirical bot-versus-human behavioral signatures, clickstream/synchronization architectures, and adversarial limitations—have been converted to knowls; introductory historical background and external news anecdotes were deliberately omitted.

References

  1. 1.Norah Abokhodair, Daisy Yoo, and David W. McDonald. 2015. Dissecting a social botnet: growth, content, and influence in Twitter. In Proceedings of the 18th ACM Conference on Computer-Supported Cooperative Work and Social Computing. ACM.
  2. 2.Luca Maria Aiello, Martina Deplano, Rossano Schifanella, and Giancarlo Ruffo. 2012. People are Strange when you’re a Stranger: Impact and Influence of Bots on Social Networks. In Proc. 6th AAAI International Conference on Weblogs and Social Media. AAAI, 10–17.
  3. 3.Lorenzo Alvisi, Allen Clement, Alessandro Epasto, Silvio Lattanzi, and Alessandro Panconesi. 2013. Sok: The evolution of sybil defense via social networks. In 2013 IEEE Symposium on Security and Privacy. IEEE, 382–396.
  4. 4.Alex Beutel, Wanhong Xu, Venkatesan Guruswami, Christopher Palow, and Christos Faloutsos. 2013. CopyCatch: stopping group attacks by spotting lockstep behavior in social networks. In Proceedings of the 22nd international conference on World Wide Web. International World Wide Web Conferences Steering Committee, 119–130.
  5. 5.Johan Bollen, Huina Mao, and Xiaojun Zeng. 2011. Twitter mood predicts the stock market. Journal of Computational Science 2, 1 (2011), 1–8.
  6. 6.Yazan Boshmaf, Ildar Muslukhov, Konstantin Beznosov, and Matei Ripeanu. 2012. Key challenges in defending against malicious socialbots. In Proceedings of the 5th USENIX Conference on Large-scale Exploits and Emergent Threats, Vol. 12.
  7. 7.Yazan Boshmaf, Ildar Muslukhov, Konstantin Beznosov, and Matei Ripeanu. 2013. Design and analysis of a social botnet. Computer Networks 57, 2 (2013), 556–578.
  8. 8.Erica J Briscoe, D Scott Appling, and Heather Hayes. 2014. Cues to Deception in Social Media Communications. In HICSS: 47th Hawaii International Conference on System Sciences. IEEE, 1435–1443.
  9. 9.Qiang Cao, Michael Sirivianos, Xiaowei Yang, and Tiago Pregueiro. 2012. Aiding the Detection of Fake Accounts in Large Scale Social Online Services. In NSDI. 197–210.
  10. 10.Qiang Cao, Xiaowei Yang, Jieqi Yu, and Christopher Palow. 2014. Uncovering Large Groups of Active Malicious Accounts in Online Social Networks. In Proceedings of the 2014 ACM SIGSAC Conference on Computer and Communications Security. ACM, 477–488.
  11. 11.Christopher A Cassa, Rumi Chunara, Kenneth Mandl, and John S Brownstein. 2013. Twitter as a sentinel in emergency situations: lessons from the Boston marathon explosions. PLoS Currents: Disasters (July 2013). DOI:http://dx.doi.org/10.1371/currents.dis.ad70cd1c8bc585e9470046cde334ee4b
  12. 12.Michael Conover, Jacob Ratkiewicz, Matthew Francisco, Bruno Gonçalves, Filippo Menczer, and Alessandro Flammini. 2011. Political polarization on Twitter. In 5th International AAAI Conference on Weblogs and Social Media. 89–96.
  13. 13.Chad Edwards, Autumn Edwards, Patric R Spence, and Ashleigh K Shelton. 2014. Is that a bot running the social media feed? Testing the differences in perceptions of communication quality for a human agent and a bot agent on Twitter. Computers in Human Behavior 33 (2014), 372–376.
  14. 14.Yuval Elovici, Michael Fire, Amir Herzberg, and Haya Shulman. 2013. Ethical considerations when employing fake identities in online social networks for research. Science and engineering ethics (2013), 1–17.
  15. 15.Aviad Elyashar, Michael Fire, Dima Kagan, and Yuval Elovici. 2013. Homing socialbots: intrusion on a specific organization’s employee using Socialbots. In Proceedings of the 2013 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining. ACM, 1358–1365.
  16. 16.Carlos A Freitas, Fabrıcio Benevenuto, Saptarshi Ghosh, and Adriano Veloso. 2014. Reverse Engineering Socialbot Infiltration Strategies in Twitter. arXiv preprint arXiv:1405.4927 (2014).
  17. 17.Rumi Ghosh, Tawan Surachawala, and Kristina Lerman. 2011. Entropy-based Classification of ‘‘Retweeting’’ Activity on Twitter. In SNA-KDD: KDD workshop on Social Network Analysis.
  18. 18.Scott A Golder and Michael W Macy. 2011. Diurnal and seasonal mood vary with work, sleep, and daylength across diverse cultures. Science 333, 6051 (2011), 1878–1881.
  19. 19.Aditi Gupta, Hemank Lamba, and Ponnurangam Kumaraguru. 2013. $1.00 per RT #BostonMarathon #PrayForBoston: Analyzing fake content on Twitter. In eCrime Researchers Summit. IEEE, 1–12.
  20. 20.Paul Heymann, Georgia Koutrika, and Hector Garcia-Molina. 2007. Fighting spam on social web sites: A survey of approaches and future challenges. Internet Computing, IEEE 11, 6 (2007), 36–45.
  21. 21.Tim Hwang, Ian Pearce, and Max Nanis. 2012. Socialbots: Voices from the fronts. Interactions 19, 2 (2012), 38–45.
  22. 22.Adam DI Kramer, Jamie E Guillory, and Jeffrey T Hancock. 2014. Experimental evidence of massive-scale emotional contagion through social networks. Proceedings of the National Academy of Sciences (2014), 201320040.
  23. 23.Kyumin Lee, Brian David Eoff, and James Caverlee. 2011. Seven Months with the Devils: A Long-Term Study of Content Polluters on Twitter. In Proceedings of the 5th International AAAI Conference on Weblogs and Social Media. 185–192.
  24. 24.Johnnatan Messias, Lucas Schmidt, Ricardo Oliveira, and Fabrıcio Benevenuto. 2013. You followed my bot! Transforming robots into influential users in Twitter. First Monday 18, 7 (2013).
  25. 25.Panagiotis T Metaxas and Eni Mustafaraj. 2012. Social media and the elections. Science 338, 6106 (2012), 472–473.
  26. 26.Abigail Paradise, Rami Puzis, and Asaf Shabtai. 2014. Anti-Reconnaissance Tools: Detecting Targeted Socialbots. Internet Computing 18, 5 (2014), 11–19.
  27. 27.Jacob Ratkiewicz, Michael Conover, Mark Meiss, Bruno Gonçalves, Alessandro Flammini, and Filippo Menczer. 2011a. Detecting and tracking political abuse in social media. In 5th International AAAI Conference on Weblogs and Social Media. 297–304.
  28. 28.Jacob Ratkiewicz, Michael Conover, Mark Meiss, Bruno Gonçalves, Snehal Patil, Alessandro Flammini, and Filippo Menczer. 2011b. Truthy: mapping the spread of astroturf in microblog streams. In WWW: 20th International Conference on World Wide Web. 249–252.
  29. 29.Tao Stein, Erdong Chen, and Karan Mangla. 2011. Facebook Immune System. In Proceedings of the 4th Workshop on Social Network Systems. ACM, 8.
  30. 30.Gianluca Stringhini, Christopher Kruegel, and Giovanni Vigna. 2010. Detecting spammers on social networks. In Proceedings of the 26th Annual Computer Security Applications Conference. ACM, 1–9.
  31. 31.Alan M Turing. 1950. Computing machinery and intelligence. Mind 49, 236 (1950), 433–460.
  32. 32.Bimal Viswanath, Ansley Post, Krishna P Gummadi, and Alan Mislove. 2011. An analysis of social network-based sybil defenses. ACM SIGCOMM Computer Communication Review 41, 4 (2011), 363–374.
  33. 33.Claudia Wagner, Silvia Mitter, Christian Korner, and Markus Strohmaier. 2012. When social bots attack:  Modeling susceptibility of users in online social networks. In WWW: 21th International Conference on World Wide Web. 41–48.
  34. 34.Randall Wald, Taghi M Khoshgoftaar, Amri Napolitano, and Chris Sumner. 2013. Predicting susceptibility to social bots on Twitter. In 2013 IEEE 14th International Conference on Information Reuse and Integration. IEEE, 6–13.
  35. 35.Gang Wang, Tristan Konolige, Christo Wilson, Xiao Wang, Haitao Zheng, and Ben Y Zhao. 2013a. You Are How You Click: Clickstream Analysis for Sybil Detection. In USENIX Security. 241–256.
  36. 36.Gang Wang, Manish Mohanlal, Christo Wilson, Xiao Wang, Miriam Metzger, Haitao Zheng, and Ben Y Zhao. 2013b. Social turing tests: Crowdsourcing sybil detection. In NDSS. The Internet Society.
  37. 37.Joseph Weizenbaum. 1966. ELIZA – a computer program for the study of natural language communication between man and machine. Commun. ACM 9, 1 (1966), 36–45.
  38. 38.Xian Wu, Ziming Feng, Wei Fan, Jing Gao, and Yong Yu. 2013. Detecting Marionette Microblog Users for Improved Information Credibility. In Machine Learning and Knowledge Discovery in Databases. Springer, 483–498.
  39. 39.Yinglian Xie, Fang Yu, Qifa Ke, Martın Abadi, Eliot Gillum, Krish Vitaldevaria, Jason Walter, Junxian Huang, and Zhuoqing Morley Mao. 2012. Innocent by association: early recognition of legitimate users. In Proceedings of the 2012 ACM conference on Computer and communications security. ACM, 353–364.
  40. 40.Zhi Yang, Christo Wilson, Xiao Wang, Tingting Gao, Ben Y Zhao, and Yafei Dai. 2014. Uncovering social network sybils in the wild. ACM Transactions on Knowledge Discovery from Data 8, 1 (2014), 2.
  41. 41.Eva Zangerle and Gunther Specht. 2014. ‘‘Sorry, I was hacked’’ A Classification of Compromised Twitter  Accounts. In SAC: the 29th Symposium On Applied Computing.

Citation

MLA
Ferrara, E., et al. “The Rise of Social Bots”. Communications of the ACM, vol. 59, no. 7, 2016, pp. 96–104, https://doi.org/10.1145/2818717.
APA
Ferrara, E., Varol, O., Davis, C., Menczer, F., & Flammini, A. (2016). The rise of social bots. Communications of the ACM, 59(7), 96–104. https://doi.org/10.1145/2818717
Chicago
Ferrara, E., O. Varol, C. Davis, F. Menczer, and A. Flammini. 2016. “The Rise of Social Bots”. Communications of the ACM 59 (7): 96–104. https://doi.org/10.1145/2818717.
Harvard
Ferrara, E. et al. (2016) “The rise of social bots”, Communications of the ACM, 59(7), pp. 96–104. Available at: https://doi.org/10.1145/2818717.
Vancouver
1. Ferrara E, Varol O, Davis C, Menczer F, Flammini A (2016) The rise of social bots. Communications of the ACM 59:96–104

BibTeX

@article{Ferrara_2016, title={The rise of social bots}, volume={59}, ISSN={1557-7317}, url={http://dx.doi.org/10.1145/2818717}, DOI={10.1145/2818717}, number={7}, journal={Communications of the ACM}, publisher={Association for Computing Machinery (ACM)}, author={Ferrara, Emilio and Varol, Onur and Davis, Clayton and Menczer, Filippo and Flammini, Alessandro}, year={2016}, month=June, pages={96–104} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF