Everyone's an influencer: quantifying influence on twitter
Eytan BakshyJake M. HofmanWinter A. MasonDuncan J. Watts
Demonstrates through an empirical analysis of 74 million information cascades that targeting large volumes of ordinary users is often more cost-effective for word-of-mouth marketing than pursuing top influencers, due to the inherent unpredictability of individual diffusion events.
Marketers and communication strategists frequently seek to accelerate word-of-mouth diffusion by identifying and seeding campaigns with influential individuals. However, conventional strategies heavily depend on retrospective accounts of viral successes rather than objective, forward-looking predictions. This creates a critical risk of misallocating resources on high-profile figures who may not reliably generate viral spread.
The article evaluates whether individual-level influence on Twitter can be predicted in advance using user attributes, past diffusion history, and content characteristics. In addition, it examines the cost-effectiveness of various influencer targeting strategies under different cost structures.
To conduct this evaluation, the study analyzed 74 million diffusion events initiated by 1.6 million seed users across an empirical Twitter follower graph consisting of approximately 56 million users and 1.7 billion connections between September and November 2009. The authors tracked information cascades of shortened web addresses from original seed posters through multi-generational reposting chains. They applied cross-validated regression tree models to forecast second-month diffusion based on first-month performance and user metrics, complemented by human content ratings on a stratified sample of web links.
The analysis yielded several key findings. First, information diffusion is extremely rare: the vast majority of shared links do not spread, resulting in an average cascade size of just 1.14 and a median of 1. Second, while large follower counts and past direct reposts are necessary precursors to reach large audiences, individual-level predictions remain relatively unreliable. Although the model accurately predicts average performance across groups, it explains only 34% of variance at the individual level. Third, evaluating content attributes—such as interestingness or emotional tone—does not improve predictive power over basic user metrics. Finally, economic simulations demonstrate that whenever the cost to identify and acquire individuals is low to moderate, targeting large numbers of ordinary users with average or below-average influence yields significantly higher reach per dollar—up to fifteen times more cost-effective than targeting top influencers.
These findings challenge the common assumption that word-of-mouth campaigns must rely on elite influencers. Because viral cascades are inherently rare and unpredictable, betting on a small number of prominent individuals introduces substantial financial risk. Instead, word-of-mouth diffusion primarily operates through many small, concurrent cascades.
Decision-makers should shift from single-influencer initiatives toward diversified, portfolio-style targeting strategies that enlist numerous ordinary users to capture reliable average effects. Prominent influencers should be reserved primarily for scenarios where high operational overhead makes managing multiple contacts cost-prohibitive. Organizations should also track objective, longitudinal diffusion metrics rather than vanity follower numbers.
These conclusions are derived from observational modeling of user-selected web links on Twitter and do not guarantee causal outcomes when content is externally sponsored. Decision-makers should treat these insights as solid, evidence-based principles for campaign design while running controlled pilot experiments to validate specific commercial applications.
- Paper: What is Twitter, a social network or a news media?, Haewoon Kwak et al. (2010). This foundational empirical study maps the topological follower graph and retweet dynamics on Twitter, establishing the baseline structural properties of user influence and information spread that the source directly builds upon.
- Paper: Maximizing the spread of influence through a social network, David Kempe et al. (2003). This paper establishes the formal algorithmic framework for influence maximization and viral seeding in social networks, providing the theoretical benchmark that the source empirically evaluates and critiques.
- Paper: Mining the network value of customers, Pedro M. Domingos et al. (2001). This work introduces the foundational marketing perspective of valuing customers by their social network influence rather than direct sales alone, motivating the source's cost-effectiveness analysis of ordinary versus top influencers.
- Paper: Cost-effective outbreak detection in networks, J. Leskovec et al. (2007). This paper establishes cost-effective sensor and node selection methods for detecting information cascades under budgetary constraints, directly informing the source's evaluation of marketing seeding strategies.
- Paper: Efficient influence maximization in social networks, Wei Chen et al. (2009). This work develops scalable heuristics for selecting influential seed nodes in large networks, providing practical context for the computational evaluation of cascade generation.
- Paper: Cheap and Fast – But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks, R. Snow et al. (2008). This paper validates the reliability of crowdsourcing subjective ratings via Amazon Mechanical Turk, providing methodological justification for the source's crowd-annotated URL characteristics.
- Paper: Information credibility on twitter, Carlos Castillo et al. (2011). Building upon findings of how content spreads across Twitter, this work evaluates the credibility of trending topics and propagation cascades using crowd-annotated assessments.
- Paper: Epidemic processes in complex networks, Romualdo Pastor-Satorras et al. (2015). This comprehensive review synthesizes mathematical models of epidemic spreading and contagion across heterogeneous complex networks, providing deeper theoretical frameworks for the empirical cascade dynamics observed in the source.
- Paper: Fake News Detection on Social Media: A Data Mining Perspective, Kai Shu et al. (2017). This survey extends the study of social media diffusion by analyzing how false information and rumors propagate through network structures.
- Paper: The COVID-19 social media infodemic, Matteo Cinelli et al. (2020). This study applies multi-platform cascade analysis to quantify the spread of online health narratives during a pandemic, extending empirical social diffusion to modern infodemic contexts.
