Everyone's an influencer: quantifying influence on twitter

Eytan BakshyJake M. HofmanWinter A. MasonDuncan J. Watts

article2011WSDM1,980 citationsWSDM 2021 Test-of-time Award

Demonstrates through an empirical analysis of 74 million information cascades that targeting large volumes of ordinary users is often more cost-effective for word-of-mouth marketing than pursuing top influencers, due to the inherent unpredictability of individual diffusion events.

Listen

Marketers and communication strategists frequently seek to accelerate word-of-mouth diffusion by identifying and seeding campaigns with influential individuals. However, conventional strategies heavily depend on retrospective accounts of viral successes rather than objective, forward-looking predictions. This creates a critical risk of misallocating resources on high-profile figures who may not reliably generate viral spread.

The article evaluates whether individual-level influence on Twitter can be predicted in advance using user attributes, past diffusion history, and content characteristics. In addition, it examines the cost-effectiveness of various influencer targeting strategies under different cost structures.

To conduct this evaluation, the study analyzed 74 million diffusion events initiated by 1.6 million seed users across an empirical Twitter follower graph consisting of approximately 56 million users and 1.7 billion connections between September and November 2009. The authors tracked information cascades of shortened web addresses from original seed posters through multi-generational reposting chains. They applied cross-validated regression tree models to forecast second-month diffusion based on first-month performance and user metrics, complemented by human content ratings on a stratified sample of web links.

The analysis yielded several key findings. First, information diffusion is extremely rare: the vast majority of shared links do not spread, resulting in an average cascade size of just 1.14 and a median of 1. Second, while large follower counts and past direct reposts are necessary precursors to reach large audiences, individual-level predictions remain relatively unreliable. Although the model accurately predicts average performance across groups, it explains only 34% of variance at the individual level. Third, evaluating content attributessuch as interestingness or emotional tonedoes not improve predictive power over basic user metrics. Finally, economic simulations demonstrate that whenever the cost to identify and acquire individuals is low to moderate, targeting large numbers of ordinary users with average or below-average influence yields significantly higher reach per dollarup to fifteen times more cost-effective than targeting top influencers.

These findings challenge the common assumption that word-of-mouth campaigns must rely on elite influencers. Because viral cascades are inherently rare and unpredictable, betting on a small number of prominent individuals introduces substantial financial risk. Instead, word-of-mouth diffusion primarily operates through many small, concurrent cascades.

Decision-makers should shift from single-influencer initiatives toward diversified, portfolio-style targeting strategies that enlist numerous ordinary users to capture reliable average effects. Prominent influencers should be reserved primarily for scenarios where high operational overhead makes managing multiple contacts cost-prohibitive. Organizations should also track objective, longitudinal diffusion metrics rather than vanity follower numbers.

These conclusions are derived from observational modeling of user-selected web links on Twitter and do not guarantee causal outcomes when content is externally sponsored. Decision-makers should treat these insights as solid, evidence-based principles for campaign design while running controlled pilot experiments to validate specific commercial applications.

  • Paper: What is Twitter, a social network or a news media?, Haewoon Kwak et al. (2010). This foundational empirical study maps the topological follower graph and retweet dynamics on Twitter, establishing the baseline structural properties of user influence and information spread that the source directly builds upon.
  • Paper: Maximizing the spread of influence through a social network, David Kempe et al. (2003). This paper establishes the formal algorithmic framework for influence maximization and viral seeding in social networks, providing the theoretical benchmark that the source empirically evaluates and critiques.
  • Paper: Mining the network value of customers, Pedro M. Domingos et al. (2001). This work introduces the foundational marketing perspective of valuing customers by their social network influence rather than direct sales alone, motivating the source's cost-effectiveness analysis of ordinary versus top influencers.
  • Paper: Cost-effective outbreak detection in networks, J. Leskovec et al. (2007). This paper establishes cost-effective sensor and node selection methods for detecting information cascades under budgetary constraints, directly informing the source's evaluation of marketing seeding strategies.
  • Paper: Efficient influence maximization in social networks, Wei Chen et al. (2009). This work develops scalable heuristics for selecting influential seed nodes in large networks, providing practical context for the computational evaluation of cascade generation.
  • Paper: Cheap and Fast – But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks, R. Snow et al. (2008). This paper validates the reliability of crowdsourcing subjective ratings via Amazon Mechanical Turk, providing methodological justification for the source's crowd-annotated URL characteristics.
  • Paper: Information credibility on twitter, Carlos Castillo et al. (2011). Building upon findings of how content spreads across Twitter, this work evaluates the credibility of trending topics and propagation cascades using crowd-annotated assessments.
  • Paper: Epidemic processes in complex networks, Romualdo Pastor-Satorras et al. (2015). This comprehensive review synthesizes mathematical models of epidemic spreading and contagion across heterogeneous complex networks, providing deeper theoretical frameworks for the empirical cascade dynamics observed in the source.
  • Paper: Fake News Detection on Social Media: A Data Mining Perspective, Kai Shu et al. (2017). This survey extends the study of social media diffusion by analyzing how false information and rumors propagate through network structures.
  • Paper: The COVID-19 social media infodemic, Matteo Cinelli et al. (2020). This study applies multi-platform cascade analysis to quantify the spread of online health narratives during a pandemic, extending empirical social diffusion to modern infodemic contexts.
Cover for Everyone's an influencer: quantifying influence on twitter

Abstract

In this paper we investigate the attributes and relative influence of 1.6M Twitter users by tracking 74 million diffusion events that took place on the Twitter follower graph over a two month interval in 2009. Unsurprisingly, we find that the largest cascades tend to be generated by users who have been influential in the past and who have a large number of followers. We also find that URLs that were rated more interesting and/or elicited more positive feelings by workers on Mechanical Turk were more likely to spread. In spite of these intuitive results, however, we find that predictions of which particular user or URL will generate large cascades are relatively unreliable. We conclude, therefore, that word-of-mouth diffusion can only be harnessed reliably by targeting large numbers of potential influencers, thereby capturing average effects. Finally, we consider a family of hypothetical marketing strategies, defined by the relative cost of identifying versus compensating potentialinfluencers.” We find that although under some circumstances, the most influential users are also the most cost-effective, under a wide range of plausible assumptions the most cost-effective performance can be realized usingordinary influencers”—individuals who exert average or even less-than-average influence.

Table of Contents

  • 1. INTRODUCTION
  • 2. RELATED WORK
  • 3. DATA
  • 4. COMPUTING INFLUENCE ON TWITTER
  • 5. PREDICTING INDIVIDUAL INFLUENCE
  • 6. THE ROLE OF CONTENT
  • 7. TARGETING STRATEGIES
  • 8. CONCLUSIONS
  • 9. ACKNOWLEDGMENTS
  • 10. REFERENCES

Knowls

  1. Knowl 1 — Cost-Effectiveness of Targeting Strategies for Seeding Network Diffusion

    empirical result

    Hypothetical word-of-mouth marketing campaigns are modeled by assigning a cost cic_i to target an individual user ii:

    ci=ca+ficfc_i = c_a + f_i c_f

    where fif_i is the number of followers of user ii, cfc_f is the marginal compensation cost per follower (benchmarked at $0.01\$0.01), and ca=αcfc_a = \alpha c_f is a fixed per-individual acquisition/management overhead cost parameterized by a multiplier α0\alpha \ge 0. The return on investment (ROI) is evaluated by measuring the influence-per-dollar, defined as the ratio of mean actual cascade size (the total number of reposts triggered by seed ii) to total targeting cost cic_i.

    Evaluating users partitioned by their predicted influence reveals:

    • When acquisition cost is zero (α=0\alpha = 0, meaning ca=$0c_a = \$0), the least influential users (who average 14\approx 14 followers and an average influence score of 0.01\approx 0.01) are the most cost-effective, generating over 1515 times greater influence-per-dollar than the highest-influence category.
    • When acquisition cost is very high (α100,000\alpha \ge 100{,}000, corresponding to ca$1,000c_a \ge \$1{,}000), the most influential individuals become the most cost-effective choice because high overhead precludes acquiring large volumes of low-follower seeds.
    • For a wide intermediate regime of acquisition costs (including α\alpha up to 10,00010{,}000, corresponding to ca=$100c_a = \$100), the most cost-effective seeds remain "ordinary influencers"—users with approximately average connectivity and influence—rather than elite super-influencers.
  2. Knowl 2 — Regression Tree Model for Predicting Information Cascade Size

    model/method

    To predict the influence of individual users on Twitter, seed-level influence is modeled using a regression tree trained via greedy recursive partitioning with 5-fold cross-validation. Individual influence is defined as log10(Sˉi)\log_{10}(\bar{S}_i), where Sˉi\bar{S}_i is the average size of all bit.ly diffusion cascades initiated by seed user ii.

    The model evaluates candidate feature sets measured during a historical observation period (Month 1):

    1. User network attributes: follower count (log10(followers+1)\log_{10}(\text{followers} + 1)), friend/following count (log10(friends+1)\log_{10}(\text{friends} + 1)), total tweet volume (log10(tweets+1)\log_{10}(\text{tweets} + 1)), and account creation date.
    2. Historical influence metrics: average, minimum, and maximum historical total cascade sizes (log10(total influence+1)\log_{10}(\text{total influence} + 1)), and average, minimum, and maximum historical local influence (log10(local influence+1)\log_{10}(\text{local influence} + 1)), where local influence denotes the number of direct reposts by immediate followers of user ii.

    During tree construction, recursive splitting selects only two features across all partitions: historical average local influence (log10(past local influence+1)\log_{10}(\text{past local influence} + 1)) and follower count (log10(followers+1)\log_{10}(\text{followers} + 1)). Historical total cascade size, out-degree (friends), tweet frequency, and account age provide no additional predictive power and are pruned from the model.

  3. Knowl 3 — Statistical Limits on Predicting Individual-Level Cascade Success

    empirical result

    When using the regression tree model to predict second-month influence from first-month performance:

    • At the aggregate partition level, the model exhibits near-perfect calibration (R2=0.98R^2 = 0.98), where the predicted mean log cascade size in each leaf node matches the actual empirical mean of that leaf node.
    • At the individual user level, unaggregated predictive accuracy is low (R2=0.34R^2 = 0.34).

    This discrepancy occurs because cascade size distributions are highly right-skewed and large cascades are exceedingly rare. Having a large follower count and strong past local influence are necessary conditions for generating large cascades, but they are far from sufficient: the vast majority of posts even by well-connected, historically successful users fail to generate cascades beyond direct followers. As a result, individual cascade sizes cannot be reliably predicted ex-ante, implying that word-of-mouth campaigns must target portfolios of seeds to realize expected average gains.

  4. Knowl 4 — Relationship Between Content Ratings and Cascade Size

    empirical result

    An analysis of 795 stratified bit.ly URLs classified and scored by Amazon Mechanical Turk human raters (with each URL receiving 3 to 20 ratings) measured attributes including:

    1. Interestingness (7-point Likert scale).
    2. Perceived interestingness to an average person (7-point Likert scale).
    3. Positive feeling / emotional valence (7-point Likert scale).
    4. Self-reported willingness to share (Email, IM, Twitter, Facebook, Digg).
    5. Site media type (Media Sharing / Social Networking, Blog / Forum, News / Mass Media, Other).
    6. Content topic category (Lifestyle, Technology, Offbeat, Entertainment, Gaming, Science, News, Business/Finance, Sports, Other).

    While bivariate analyses indicate that content rated as more interesting or evoking more positive emotion produces larger average cascade sizes, and media-sharing/lifestyle content out-diffuses news content on average, incorporating these content features into the regression tree does not improve predictive accuracy (R2=0.31R^2 = 0.31 compared to R2=0.34R^2 = 0.34 for the user-only model). Because non-successful events overwhelmingly dominate the population, content attributes fail to reliably differentiate successful from unsuccessful cascades at the level of individual diffusion events.

  5. Knowl 5 — Seed-Initiated Information Cascade and Influence Metric

    definition

    On the Twitter follower graph, an information diffusion cascade is defined as a directed tree of repostings of a specific URL initiated by a "seed" user. A user is designated as a seed for a URL if they post the URL without having previously received it from any user they follow on the follower graph.

    The influence of a seed user ii for a given URL post is quantified by the cascade size SS, defined as the total number of nodes in the influence tree (including the seed and all subsequent reposting descendants reachable via follower links). The user's overall influence score is defined as the mean cascade size across all URLs seeded by that user during the measurement period.

  6. Knowl 6 — Attribution Rules for Influence with Multiple Prior-Posting Friends

    model/method

    When a user BB posts a URL after k>1k > 1 of their friends (users whom BB follows) have already posted the same URL, credit for influencing BB can be attributed using three distinct rules:

    1. First Influence: Full influence credit (1.01.0) is assigned to the friend who posted the URL earliest. This assumes information adoption is driven by initial exposure/primacy.
    2. Last Influence: Full influence credit (1.01.0) is assigned to the friend who posted the URL most recently prior to BB's post. This assumes adoption is triggered by the most recent exposure.
    3. Split Influence: Influence credit is divided equally among all kk prior-posting friends, each receiving a weight of 1/k1/k. This assumes the propensity to adopt accumulates monotonically with multiple exposures.

    Constructing disjoint influence trees under all three rules yields slightly differing numerical values but identical qualitative conclusions regarding cascade structures and predictive modeling.

  7. Knowl 7 — Heavy-Tailed Distributions of Cascade Size and Depth on Twitter

    empirical result

    Across 74 million bit.ly URL diffusion events tracked over a two-month period on Twitter:

    • Cascade Size: The distribution of cascade sizes follows an approximate power-law distribution. The median cascade size is 11 and the mean is 1.141.14, indicating that the overwhelming majority of seed posts never propagate beyond the original poster, while a minute fraction spread to thousands of reposts.
    • Cascade Depth: The distribution of cascade depth (the length of the longest directed path from the seed to a reposting leaf) is right-skewed and approximately exponential. The median depth is 00 (only the seed node), with rare maximal cascades reaching up to 99 generations of propagation. Non-trivial cascades that do spread predominantly exhibit a depth of 11.
  8. Knowl 8 — Summary Statistics of the Active bit.ly Twitter Follower Graph and Seeding Dataset

    data/table

    Tracking all public tweets broadcast on Twitter between September 13, 2009, and November 15, 2009 (excluding an API outage during October 14–16) yielded 87 million bit.ly URL diffusion events from 1.03 billion total tweets. Restricting the set to seed users active in both consecutive 30-day observation months resulted in 74 million diffusion events across 1.6 million seed users. The corresponding follower subgraph crawling yielded 56 million users and 1.7 billion directed edges.

    Metric # Followers # Friends # Seeds Posted
    Median 85.00 82.00 11.00
    Mean 557.10 294.10 46.33
    Max 3,984,000.00 759,700.00 54,890

    The table summarizes the in-degree (followers), out-degree (friends), and seeding activity per user in the active bit.ly crawling sample. The distribution of followers is more skewed than friends (maximum follower count near 4M vs. 760K friends), reflecting the asymmetric, one-way nature of Twitter follow relationships and active-user sampling bias.

  9. Knowl 9 — Observational and Behavioral Limitations in Measuring Social Influence

    limitation

    The empirical estimation and predictive modeling of Twitter diffusion cascades have several key limitations:

    1. Lack of Causal Attribution: Findings are derived from observational data where users voluntarily choose content to post. Exogenous seeding by marketers (e.g., sponsored posts) may alter user behavior, reception, and diffusion dynamics compared to organic sharing.
    2. Homophily vs. Influence Confounding: Defining influence purely through temporal follower links may capture homophily (connected users sharing similar interests independently posting identical URLs) rather than true causal influence, making observed cascade sizes an upper bound.
    3. Narrow Influence Definition: Cascade size strictly measures the mechanical transmission of URLs rather than offline behavioral change, purchasing decisions, or shifts in political beliefs.
    4. Unobserved Consumption: Reposting represents an active, high-threshold signal of influence; it ignores passive consumption (click-throughs and reading), which cannot be reliably separated from automated web crawler hits.

Coverage note — No substantial contributed material was omitted.

References

  1. 1.E. Adar and A. Adamic, Lada. Tracking information epidemics in blogspace. In 2005 IEEE/WIC/ACM International Conference on Web Intelligence, Compiegne University of Technology, France, 2005.
  2. 2.S. Aral, L. Muchnik, and A. Sundararajan. Distinguishing influence-based contagion from homophily-driven diffusion in dynamic networks. Proceedings of the National Academy of Sciences, 106(51):21544, 2009.
  3. 3.E. Bakshy, B. Karrer, and A. Adamic, Lada. Social influence and the diffusion of user-created content. In 10th ACM Conference on Electronic Commerce, Stanford, California, 2009. Association of Computing Machinery.
  4. 4.F. M. Bass. A new product growth for model consumer durables. Management Science, 15(5):215–227, 1969.
  5. 5.R. A. Berk. An introduction to sample selection bias in sociological data. American Sociological Review, 48(3):386–398, 1983.
  6. 6.L. Breiman, J. Friedman, R. Olshen, and C. Stone. Classification and regression trees. Chapman & Hall/CRC, 1984.
  7. 7.M. Cha, H. Haddadi, F. Benevenuto, and K. P. Gummad. Measuring user influence on twitter: The million follower fallacy. In 4th Int'l AAAI Conference on Weblogs and Social Media, Washington, DC, 2010.
  8. 8.R. M. Dawes. Everyday irrationality: How pseudo-scientists, lunatics, and the rest of us systematically fail to think rationally. Westview Pr, 2002.
  9. 9.J. Denrell. Vicarious learning, undersampling of failure, and the myths of management. Organization Science, 14(3):227–243, 2003.
  10. 10.M. Gladwell. The Tipping Point: How Little Things Can Make a Big Difference. Little Brown, New York, 2000.
  11. 11.J. Goldenberg, S. Han, D. R. Lehmann, and J. W. Hong. The role of hubs in the adoption process. Journal of Marketing, 73(2):1–13, 2009.
  12. 12.A. Goyal, F. Bonchi, and L. V. S. Lakshmanan. Discovering leaders from community actions. pages 499–508. ACM, 2008. Proceeding of the 17th ACM conference on Information and knowledge management.
  13. 13.D. Gruhl, R. Guha, D. Liben-Nowell, and A. Tomkins. Information diffusion through blogspace. pages 491–501. ACM New York, NY, USA, 2004.
  14. 14.E. Katz and P. F. Lazarsfeld. Personal influence; the part played by people in the flow of mass communications. Free Press, Glencoe, Ill.” 1955.
  15. 15.E. Keller and J. Berry. The Influentials: One American in Ten Tells the Other Nine How to Vote, Where to Eat, and What to Buy. Free Press, New York, NY, 2003.
  16. 16.D. Kempe, J. Kleinberg, and E. Tardos. Maximizing the spread of influence through a social network. In 9th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Washington, DC, USA., 2003. Association of Computing Machinery.
  17. 17.A. Kittur, E. H. Chi, and B. Suh. Crowdsourcing user studies with mechanical turk. Proceedings of the twenty-sixth annual SIGCHI conference on Human factors in computing systems, 2008.
  18. 18.H. Kwak, C. Lee, H. Park, and S. Moon. What is twitter, a social network or a news media? pages 591–600. ACM, 2010.
  19. 19.A. Leavitt, E. Burchard, D. Fisher, and S. Gilbert. The influentials: New approaches for analyzing influence on twitter.
  20. 20.J. Leskovec, A. Adamic, Lada, and A. Huberman, Bernardo. The dynamics of viral marketing. ACM Trans. Web, 1(1):5, 2007.
  21. 21.D. Liben-Nowell and J. Kleinberg. Tracing information flow on a global scale using internet chain-letter data. Proceedings of the National Academy of Sciences, 105(12):4633, 2008.
  22. 22.W. Mason and S. Suri. Conducting Behavioral Research on Amazon’s Mechanical Turk. SSRN eLibrary, 2010.
  23. 23.W. Mason and D. J. Watts. Financial incentives and the performance of crowds. Proceedings of the ACM SIGKDD Workshop on Human Computation, pages 77–85, 2009.
  24. 24.S. A. Munson and P. Resnick. Presenting diverse political opinions: how and how much. pages 1457–1466. ACM, 2010.
  25. 25.R. R. Picard and R. D. Cook. Cross-validation of regression models. Journal of the American Statistical Association, 79(387):575–583, 1984.
  26. 26.E. M. Rogers. Diffusion of innovations. Free Press, New York, 4th edition, 1995.
  27. 27.V. S. Sheng, F. Provost, and P. G. Ipeirotis. Get another label? improving data quality and data mining using multiple, noisy labelers. pages 614–622. ACM, 2008.
  28. 28.R. Snow, B. O’Connor, D. Jurafsky, and A. Y. Ng. Cheap and fast-but is it good? evaluating non-expert annotations for natural language tasks, 2008.
  29. 29.E. S. Sun, I. Rosenn, C. A. Marlow, and T. M. Lento. Gesundheit! modeling contagion through facebook news feed. In International Conference on Weblogs and Social Media, San Jose, CA, 2009. AAAI.
  30. 30.B. Tomlinson and C. Cockram. Sars: Experience at prince of wales hospital, hong kong. The Lancet, 361(9368):1486–1487, 2003.
  31. 31.D. J. Watts. A simple model of information cascades on random networks. Proceedings of the National Academy of Science, U.S.A., 99:5766–5771, 2002.
  32. 32.D. J. Watts and P. S. Dodds. Influentials, networks, and public opinion formation. Journal of Consumer Research, 34:441–458, 2007.
  33. 33.D. J. Watts and J. Peretti. Viral marketing for the real world. Harvard Business Review, May:22–23, 2007.
  34. 34.G. Weimann. The Influentials: People Who Influence People. State University of New York Press, Albany, NY, 1994.
  35. 35.J. Weng, E. P. Lim, J. Jiang, and Q. He. Twitterrank: finding topic-sensitive influential twitterers. pages 261–270. ACM, 2010.

Citation

MLA
Bakshy, E., et al. “Everyone's an Influencer”. Proceedings of the Fourth ACM International Conference on Web Search and Data Mining, 2011, pp. 65–74, https://doi.org/10.1145/1935826.1935845.
APA
Bakshy, E., Hofman, J. M., Mason, W. A., & Watts, D. J. (2011). Everyone's an influencer. Proceedings of the Fourth ACM International Conference on Web Search and Data Mining, 65–74. https://doi.org/10.1145/1935826.1935845
Chicago
Bakshy, E., J. M. Hofman, W. A. Mason, and D. J. Watts. 2011. “Everyone's an Influencer”. Proceedings of the Fourth ACM International Conference on Web Search and Data Mining, 65–74. https://doi.org/10.1145/1935826.1935845.
Harvard
Bakshy, E. et al. (2011) “Everyone's an influencer”, Proceedings of the fourth ACM international conference on Web search and data mining. ACM, pp. 65–74. Available at: https://doi.org/10.1145/1935826.1935845.
Vancouver
1. Bakshy E, Hofman JM, Mason WA, Watts DJ (2011) Everyone's an influencer. In: Proceedings of the fourth ACM international conference on Web search and data mining. ACM, pp 65–74

BibTeX

@inproceedings{Bakshy_2011, series={WSDM′11}, title={Everyone’s an influencer: quantifying influence on twitter}, url={http://dx.doi.org/10.1145/1935826.1935845}, DOI={10.1145/1935826.1935845}, booktitle={Proceedings of the fourth ACM international conference on Web search and data mining}, publisher={ACM}, author={Bakshy, Eytan and Hofman, Jake M. and Mason, Winter A. and Watts, Duncan J.}, year={2011}, month=Feb, pages={65–74}, collection={WSDM′11} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF