Bridging Prediction and Intervention Problems in Social Systems
Lydia T. LiuInioluwa Deborah RajiAngela ZhouLuke GuerdanJessica HullmanDaniel MalinskyBryan WilderSimone ZhangHammaad AdamAmanda Coston
Proposes a paradigm shift that models automated decision systems as policy interventions rather than standalone prediction tools, unifying statistical methods to study their real-world design, implementation, and societal outcomes.
Automated decision systems (ADS) are widely deployed across high-stakes social sectors—including criminal justice, healthcare, education, social welfare, and finance—to forecast individual risks and allocate institutional actions. However, standard development practices treat these tools as isolated statistical prediction tasks. In practice, systems that achieve high predictive accuracy on historical benchmarks frequently fail to improve real-world human outcomes or even actively exacerbate systemic harms. The article addresses this critical gap between algorithmic prediction and actual policy effectiveness, evaluating why this disconnect occurs and demonstrating how organizations must transition from a narrow predictive mindset to an interventional paradigm that accounts for downstream decisions, causal effects, and operational contexts.
The article demonstrates that deploying an automated prediction model is inherently a policy intervention within a complex sociotechnical system. By evaluating historical case studies and modern causal inference frameworks, the analysis shows that optimizing a statistical score does not guarantee optimal decision-making. First, prioritizing individuals based purely on baseline prognostic risk (such as predicted recidivism or future medical costs) often fails because baseline risk does not inherently correlate with treatment efficacy (how much an individual actually benefits from an intervention). Second, standard predictive benchmarks suffer from pervasive validity failures, notably label bias (where available proxies such as arrests or billing costs misrepresent true underlying constructs like criminal activity or health need) and selective label bias (where past institutional decisions censor counterfactual outcomes). Third, empirical evaluations demonstrate that human decision-makers introduce behavioral discretion, inconsistencies, and unobserved confounding that standard predictive models fail to accommodate.
These findings have critical implications for organizational risk, compliance, and resource allocation. Deploying predictive tools without causal grounding risks wasting public and private investments on ineffective interventions, perpetuating institutional bias, and facing legal vulnerabilities regarding procedural due process. Moreover, the results challenge the standard assumption that complex machine learning models inherently outperform simple heuristics in social contexts. In highly resource-constrained environments, the analysis highlights that investing directly in expanded service capacity often yields substantially greater social welfare gains than marginal improvements in predictive targeting accuracy.
To move forward, the article outlines four strategic paths for leadership and practitioners. First, institutions must critically interrogate whether an algorithmic solution is truly warranted, explicitly comparing proposed tools against non-algorithmic alternatives such as bureaucratic rules, lotteries, or expanded base resources. Second, when automated support is justified, technical teams should re-engineer problem formulations to target causal treatment effects (such as individualized treatment rules or counterfactual risk models) rather than isolated predictions. Third, organizations should establish robust interventional evaluation frameworks—leveraging randomized rollouts, quasi-experimental designs, and human-AI interaction modeling—to measure end-to-end downstream impacts. Finally, leaders must implement sustainable governance and auditing structures that incorporate participatory stakeholder feedback, monitor total cost of ownership, and ensure meaningful procedural recourse for affected individuals.
Decision-makers must maintain caution regarding these findings due to the inherent difficulty of identifying unobserved confounding in retrospective administrative data and the presence of unavoidable trade-offs between competing social objectives. Causal conclusions remain conditional on specific structural assumptions, and randomized evaluations can be limited by organizational constraints. Nevertheless, there is high confidence in the foundational finding: assessing automated systems solely on static predictive accuracy provides insufficient evidence of real-world benefit, necessitating an interventional approach to ADS design, deployment, and governance.
- Paper: Fairness and Abstraction in Sociotechnical Systems, Andrew D. Selbst et al. (2019). Selbst et al. show why fairness analysis must account for the surrounding sociotechnical system, a foundation for this paper’s shift from isolated prediction to intervention.
- Paper: Fairness Without Demographics in Repeated Loss Minimization, Tatsunori B. Hashimoto et al. (2018). This paper explains how repeated model deployment can reshape participation and worsen group outcomes, clarifying the feedback dynamics behind an intervention-oriented view.
- Paper: On the Fairness of Causal Algorithmic Recourse, Julius von Kügelgen et al. (2022). Its causal account of how recourse actions change downstream outcomes supplies a concrete example of why social-system interventions cannot be reduced to prediction.
- Paper: Generalization Bounds and Representation Learning for Estimation of Potential Outcomes and Causal Effects, Fredrik D. Johansson et al. (2022). Its framework for estimating potential outcomes under interventions provides the causal-inference tools needed to understand the paper’s broader decision-and-outcome setup.
- Paper: Can Revealed Preferences Clarify LLM Alignment and Steering?, Khurram Yamin et al. (2026). It carries the paper’s focus on decisions beyond predictions into a framework for testing whether language models’ choices reflect coherent preferences and respond reliably to steering.
- Paper: MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning, Haohan Yu et al. (2026). It advances an intervention-based account of multi-agent systems by using counterfactual actions to measure how one agent’s choices causally affect teammates and shared outcomes.
