Bridging Prediction and Intervention Problems in Social Systems

Lydia T. LiuInioluwa Deborah RajiAngela ZhouLuke GuerdanJessica HullmanDaniel MalinskyBryan WilderSimone ZhangHammaad AdamAmanda Coston

article2025arXiv13 citationsBest Paper Award at the RSS 2025 Generative Models x HRI

Proposes a paradigm shift that models automated decision systems as policy interventions rather than standalone prediction tools, unifying statistical methods to study their real-world design, implementation, and societal outcomes.

Listen

Automated decision systems (ADS) are widely deployed across high-stakes social sectors—including criminal justice, healthcare, education, social welfare, and finance—to forecast individual risks and allocate institutional actions. However, standard development practices treat these tools as isolated statistical prediction tasks. In practice, systems that achieve high predictive accuracy on historical benchmarks frequently fail to improve real-world human outcomes or even actively exacerbate systemic harms. The article addresses this critical gap between algorithmic prediction and actual policy effectiveness, evaluating why this disconnect occurs and demonstrating how organizations must transition from a narrow predictive mindset to an interventional paradigm that accounts for downstream decisions, causal effects, and operational contexts.

The article demonstrates that deploying an automated prediction model is inherently a policy intervention within a complex sociotechnical system. By evaluating historical case studies and modern causal inference frameworks, the analysis shows that optimizing a statistical score does not guarantee optimal decision-making. First, prioritizing individuals based purely on baseline prognostic risk (such as predicted recidivism or future medical costs) often fails because baseline risk does not inherently correlate with treatment efficacy (how much an individual actually benefits from an intervention). Second, standard predictive benchmarks suffer from pervasive validity failures, notably label bias (where available proxies such as arrests or billing costs misrepresent true underlying constructs like criminal activity or health need) and selective label bias (where past institutional decisions censor counterfactual outcomes). Third, empirical evaluations demonstrate that human decision-makers introduce behavioral discretion, inconsistencies, and unobserved confounding that standard predictive models fail to accommodate.

These findings have critical implications for organizational risk, compliance, and resource allocation. Deploying predictive tools without causal grounding risks wasting public and private investments on ineffective interventions, perpetuating institutional bias, and facing legal vulnerabilities regarding procedural due process. Moreover, the results challenge the standard assumption that complex machine learning models inherently outperform simple heuristics in social contexts. In highly resource-constrained environments, the analysis highlights that investing directly in expanded service capacity often yields substantially greater social welfare gains than marginal improvements in predictive targeting accuracy.

To move forward, the article outlines four strategic paths for leadership and practitioners. First, institutions must critically interrogate whether an algorithmic solution is truly warranted, explicitly comparing proposed tools against non-algorithmic alternatives such as bureaucratic rules, lotteries, or expanded base resources. Second, when automated support is justified, technical teams should re-engineer problem formulations to target causal treatment effects (such as individualized treatment rules or counterfactual risk models) rather than isolated predictions. Third, organizations should establish robust interventional evaluation frameworks—leveraging randomized rollouts, quasi-experimental designs, and human-AI interaction modeling—to measure end-to-end downstream impacts. Finally, leaders must implement sustainable governance and auditing structures that incorporate participatory stakeholder feedback, monitor total cost of ownership, and ensure meaningful procedural recourse for affected individuals.

Decision-makers must maintain caution regarding these findings due to the inherent difficulty of identifying unobserved confounding in retrospective administrative data and the presence of unavoidable trade-offs between competing social objectives. Causal conclusions remain conditional on specific structural assumptions, and randomized evaluations can be limited by organizational constraints. Nevertheless, there is high confidence in the foundational finding: assessing automated systems solely on static predictive accuracy provides insufficient evidence of real-world benefit, necessitating an interventional approach to ADS design, deployment, and governance.

arXiv: 2507.05216
Cover for Bridging Prediction and Intervention Problems in Social Systems

Abstract

Many automated decision systems (ADS) are designed to solve prediction problems -- where the goal is to learn patterns from a sample of the population and apply them to individuals from the same population. In reality, these prediction systems operationalize holistic policy interventions in deployment. Once deployed, ADS can shape impacted population outcomes through an effective policy change in how decision-makers operate, while also being defined by past and present interactions between stakeholders and the limitations of existing organizational, as well as societal, infrastructure and context. In this work, we consider the ways in which we must shift from a prediction-focused paradigm to an intervention-oriented paradigm when considering the impact of ADS within social systems. We argue this requires a new default problem setup for ADS beyond prediction, to instead consider predictions as decision support, final decisions, and outcomes. We highlight how this perspective unifies modern statistical frameworks and other tools to study the design, implementation, and evaluation of ADS systems, and point to the research directions necessary to operationalize this paradigm shift. Using these tools, we characterize the limitations of focusing on isolated prediction tasks, and lay the foundation for a more intervention-oriented approach to developing and deploying ADS.

Table of Contents

  • 1 Introduction
  • Scope
  • 2 Background
  • Details of the Institutional Context
  • 3 Model Design
  • Setup: Predictions and Interventions
  • Predictive vs. Interventional targeting
  • Decision-theoretic Formulations
  • Reconsidering Problem Formulation
  • Ecosystem of Problem Formulation
  • Domain and empirical knowledge is needed to justify data-driven targeting
  • Causal and Interventional Analysis Can Guide Predictive Task Design
  • 4 Evaluation Science
  • Benchmarking Failures for ADS
  • Validity
  • Label Bias
  • Selection bias from prior decisions
  • Towards an Interventional Evaluation Paradigm
  • Evaluating ITR performance is a problem of off-policy evaluation
  • Evaluating ADS with quasi-experimental and observational causal analysis
  • Evaluating human-decision-makers’ use of ADS
  • Experimental and in-deployment evaluations
  • Broader Social Challenges in Evaluating ADS
  • Defining “Better” Decision-Making in Evaluations
  • Institutional Inter-dependencies, Spillover Effects, and Bounding the Scope of ADS Evaluation
  • Measurement Challenges and Goodhart’s Law
  • 5 Implementation Science
  • Individual Deployment Context
  • Interaction Design
  • Intervention Design
  • Institutional Deployment Context
  • Integration & Change Management
  • Implementation Factors & Considerations
  • Societal Context
  • 6 Prescriptive Angles
  • Path 1: Interrogating the Role of Prediction
  • Path 2: Better and More Engineering
  • Path 3: Understanding and assessing models in deployment
  • Path 4: Sustainable Governance and Maintenance of ADS
  • A Additional discussion on Background
  • Additional examples
  • B Additional discussion on Model Design
  • More detail on WPRS
  • Technical formulations: Decision-aware learning
  • Additional discussion on AFST
  • Implications of embeddedness of decisions in the social world on problem formulations
  • Spurious correlations
  • C Additional discussion on Evaluation Science
  • Challenges and Open Problems for RCTs of ADS in Social Settings
  • D Additional discussion on Implementation Science
  • References

Knowls

  1. Knowl 1 — Automated decision systems should be formulated as interventions in institutional processes

    model/method

    An automated decision system (ADS) in a consequential social setting is not just a predictor: it is one component of an institutional process that connects information about people to decisions and downstream outcomes. A useful default setup distinguishes pre-decision covariates or context XX, institutional decisions or actions DD, and outcomes of interest YY. A prediction RR can affect outcomes by influencing a decision-maker’s runtime decision, while the institution’s resources, goals, constraints, and existing practices also shape the process. The relevant comparison is therefore not only between predictive models, but between the ADS-assisted process and its bureaucratic counterfactual: the organization’s existing decision process without the ADS, which may involve categorical rules and human discretion. This framing makes institutional and policy context part of the problem definition, rather than treating deployment as a neutral step after model development.

  2. Knowl 2 — Risk targeting and treatment-effect targeting answer different allocation questions

    model/method

    For a binary intervention D∈{0,1}D\in\{0,1\}, let Y(0)Y(0) and Y(1)Y(1) be an individual’s potential outcomes without and with the intervention, and let X=xX=x denote that individual’s covariates. Baseline-risk targeting prioritizes people with high expected outcomes in the absence of intervention, whereas individualized treatment targeting prioritizes people whose outcomes are expected to improve most under intervention. For adverse outcomes, the baseline-risk rule can be written as πrisk(x)=1{E[Y(0)∣X=x]>t}\pi_{\mathrm{risk}}(x)=\mathbf{1}\{E[Y(0)\mid X=x]>t\}; for beneficial outcomes, a treatment-effect rule is πITR(x)=1{E[Y(1)−Y(0)∣X=x]>t}\pi_{\mathrm{ITR}}(x)=\mathbf{1}\{E[Y(1)-Y(0)\mid X=x]>t\}. Here π\pi is the allocation rule and tt is a threshold; under a capacity constraint, the threshold can be chosen to allocate the available budget to the highest-ranked treatment effects. The rules need not select the same people: high baseline risk does not guarantee high intervention benefit. Baseline-risk allocation may nevertheless express a deliberate priority for the worst off, so the choice requires making the institution’s goals and equity-efficiency tradeoffs explicit. Models of outcomes under each intervention, E[Y(d)∣X]E[Y(d)\mid X], can also make intervention scenarios explicit; a model of pooled historical outcomes E[Y∣X]E[Y\mid X] may hide variation across past decisions and be vulnerable when the intervention mix changes.

  3. Knowl 3 — Prediction is sufficient only when the decision loss is determined by observable outcomes

    model/method

    Whether a social decision problem can be solved by prediction alone depends on how decisions relate to the outcome and objective. If the decision does not change the outcome being predicted and the institution can specify a loss ℓ(D,Y)\ell(D,Y) from observable decision-outcome pairs, historical data can support evaluating the loss associated with alternative decisions. If the decision changes the outcome and the objective depends on the result under that decision, the relevant quantity is instead Y(D)Y(D); evaluating an alternative decision then requires counterfactual outcomes that are not jointly observed. In that case, targeting generally requires causal reasoning, such as estimating treatment effects, under assumptions sufficient to identify them from the available data. The paper also distinguishes plug-in prediction, direct optimization over decision rules, and decision-aware learning, which chooses a predictor based on the downstream loss induced when its predictions are used in a decision procedure. A cited sufficient-action result gives a further criterion: prediction can be adequate for targeting when an action improves outcomes for all decision subjects regardless of their individual circumstances.

  4. Knowl 4 — Evaluate ADS along the causal chain from prediction to decision to outcome

    model/method

    Predictive benchmark metrics assess how a score RR predicts a chosen outcome, but do not by themselves show whether the ADS changes institutional decisions or improves outcomes. The paper organizes evaluation around the chain X→R→D→YX\rightarrow R\rightarrow D\rightarrow Y: studies of D→YD\rightarrow Y assess interventions or treatment rules; policy-level before-after or quasi-experimental studies assess aggregate effects of introducing an ADS; studies of R→DR\rightarrow D examine how decision-makers respond to scores; and individual-level studies can analyze the linked prediction, decision, and outcome process. The appropriate design depends on whether assignment of ADS access or decisions can be randomized and whether individual-level data or only aggregate outcomes are available. Randomized trials can provide strong causal evidence when feasible, while observational and quasi-experimental methods can contribute evidence when randomization is not possible. Evaluation may need both pre-deployment analyses and post-deployment studies, because score performance alone does not establish deployment impact.

  5. Knowl 5 — Proxy outcomes and decision-censored data can invalidate ADS evaluations

    limitation

    Evaluations can fail to measure the intended construct even when their calculations are correct. A proxy outcome may not represent the policy goal: healthcare costs, for example, can understate the health needs of Black patients relative to white patients with similar needs because access to care differs; re-arrest can also be an imperfect measure of actual criminal activity because enforcement and reporting affect the label. Evaluations can also use a selected sample whose outcomes were made observable by prior decisions. Detained defendants cannot experience failure to appear while detained, and lenders do not observe repayment outcomes for applicants they reject; performance calculated on released defendants or approved borrowers therefore need not represent performance for the full population. These problems can distort both measurement and generalization. Potential responses include collecting more direct outcomes or representative samples when feasible, evaluating multiple plausible proxies, using sensitivity analysis for proxy-target relationships, and applying methods that address measurement error or selectively missing outcomes.

  6. Knowl 6 — Cook County example shows why risk-based supervision may not maximize intervention benefit

    empirical result

    The paper uses Cook County pretrial data to illustrate the difference between allocating supervised release by baseline risk and allocating it by estimated treatment effect. The county’s risk-based recommendation can be represented as recommending supervision when the predicted risks of failure to appear (FTA) and new criminal activity (NCA) without supervision, summed conditional on covariates XX, exceed a threshold tt. The authors estimate average treatment effects (ATEs), with 90% confidence intervals, across FTA-risk levels and racial groups, and examine estimated heterogeneous treatment-effect distributions. The estimated ATE for FTA is weakly decreasing across higher risk levels, while the estimated treatment effects are larger on average for white non-Hispanic defendants than for Black and white-Hispanic defendants. Thus, a risk threshold can identify high-risk people and may also identify people who benefit, but differences in the magnitude of estimated benefits across groups can make risk-based allocation less targeting-efficient than a threshold on treatment effects. The analysis is exploratory, excludes an Other racial category because of smaller sample size, and is not presented as sufficient evidence for a policy prescription; the authors stress the need to examine treatment-effect heterogeneity, mechanisms, and normative priorities.

  7. Knowl 7 — Randomized PSA access had little overall effect on judges’ decisions in Dane County

    empirical result

    In a 30-month Dane County, Wisconsin study spanning 2017–2019, judges were randomly assigned whether to see a Public Safety Assessment (PSA) score for each pretrial case; the score was calculated for every case. Researchers tracked judges’ decisions among signature bonds and small or large bail, as well as defendant outcomes including failure to appear, new criminal activity, and new violent criminal activity over the following two years. Making the PSA visible had a negligible overall effect on judges’ decisions and produced few notable differences in downstream outcomes. Effects differed by gender: visible PSA scores led to more lenient decisions for female defendants, while male defendants were relatively unaffected. The study did not find a significant disparate impact of PSA visibility by race. The example illustrates that providing a prediction to a human decision-maker does not necessarily change decisions or outcomes, and that average effects can conceal subgroup differences.

  8. Knowl 8 — Human-AI evaluation must account for information access and who retains decision control

    model/method

    The paper organizes human-AI interaction design around two questions: whether the human and model have access to the same or different information, and whether a human must retain control of the final decision. If information is similar and human control is not required, automation may be appropriate; if human control is required, interpretability and prediction-uncertainty information are candidate supports. If information differs and the model can decide, learning-to-defer methods can allocate cases between human and model; if the human retains control, uncertainty communication should account for both agents’ information, for example through multi-calibrated uncertainty. Evaluation of model reliance should not simply count acceptance of recommendations or ex-post switches judged against realized outcomes: probabilistic outcomes mean an apparently incorrect choice may have been reasonable given information available at decision time. A decision-theoretic alternative compares observed behavior with an idealized Bayesian decision-maker who knows the relevant signal and outcome distributions, and can estimate whether access to both human and model judgments could outperform either alone.

  9. Knowl 9 — Implementation context shapes whether an ADS changes decisions and outcomes

    model/method

    Implementation science is the systematic study of methods and strategies that support integration and sustained use of evidence-based ADS in real-world settings. The paper treats deployment context as operating at multiple scales: an individual interaction, the intervention made available in response to a prediction, the institution’s workflows and resources, and the broader societal and regulatory environment. Relevant individual-level factors include how a score is displayed, when it is shown, the time available to decide, human discretion, and whether an effective response or sufficient service capacity exists. Institutional deployment can involve multiple decisions, actors, and prediction systems rather than a single score and decision; change management must account for stakeholders whose incentives may not align with those of decision subjects. Stakes, recourse, capacity constraints, timing, and legal compatibility can all alter how otherwise identical models function in practice. Accordingly, understanding deployment requires documenting and assessing the context, not merely reporting model accuracy.

  10. Knowl 10 — The paper recommends four complementary paths beyond accuracy-focused model development

    model/method

    The paper identifies four directions for making ADS more responsive to social outcomes. First, interrogate whether prediction is needed at all: alternatives include expanding resources, improving interventions, or using lotteries for allocation. Second, improve engineering by choosing an appropriate estimand and decision formulation, revisiting data collection, and incorporating justified constraints or desiderata. Third, assess systems in deployment through rigorous experimental designs together with participatory and qualitative approaches that surface efficacy, safety, procedural effects, and harms to affected communities. Fourth, establish sustainable governance and maintenance through institutional accountability, ongoing performance monitoring, cost-benefit assessment, codified practices, and regulatory engagement. These directions are proposals for research and practice, not claims that any one approach is universally appropriate.

Coverage note — Additional domain examples and case studies—including housing, unemployment services, healthcare, education, and the Kentucky pretrial event study—are omitted as separate knowls because they mainly illustrate the model-design, evaluation, and implementation principles captured here rather than adding distinct framework-level contributions.

References

  1. 1.How New Jersey Used an Algorithm To Drastically Reduce Its Jail Population – And Why It Might Not Be the Right Tool for the Job — aclu-nj.org. https://www.aclu-nj.org/en/news/how-new-jersey-used-algorithm-drastically-reduce-its-jail-population-and-why-it-might-not-be. [Accessed 02-11-2024].
  2. 2.R. Abebe, S. Barocas, J. Kleinberg, K. Levy, M. Raghavan, and D. G. Robinson. Roles for computing in social change. In Proceedings of the 2020 conference on fairness, accountability, and transparency, pages 252–260, 2020.
  3. 3.E. Achterhold, M. Mühlböck, N. Steiber, and C. Kern. Fairness in algorithmic profiling: The amas case. Minds and Machines, 35(1):9, 2025.
  4. 4.ACLU Idaho. acluidaho.org. https://www.acluidaho.org/sites/default/files/field_documents/aclu_pess_release_judge_rules_against_idaho_medicaid_program_2023-09-07.pdf. [Accessed 29-03-2025].
  5. 5.S. Ægisdóttir, M. J. White, P. M. Spengler, A. S. Maugherman, L. A. Anderson, R. S. Cook, C. N. Nichols, G. K. Lampropoulos, B. S. Walker, G. Cohen, et al. The meta-analysis of clinical judgment project: Fifty-six years of accumulated research on clinical versus statistical prediction. The counseling psychologist, 34(3):341–382, 2006.
  6. 6.A. Albright. If you give a judge a risk score: evidence from kentucky bail decisions. Law, Economics, and Business Fellows’ Discussion Paper Series, 85, 2019.
  7. 7.A. Alkhatib and M. Bernstein. Street-level algorithms: A theory at the gaps between policy and decisions. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pages 1–13, 2019.
  8. 8.Allegheny County Analytics. alleghenycountyanalytics.us. https://www.alleghenycountyanalytics.us/wp-content/uploads/2018/10/17-ACDHS-11_AFST_102518.pdf#page=5.44. [Accessed 10-11-2024].
  9. 9.R. Alur, L. Laine, D. Li, M. Raghavan, D. Shah, and D. Shung. Auditing for human expertise. Advances in Neural Information Processing Systems, 36, 2024.
  10. 10.D. Ames, C. Handan-Nader, D. E. Ho, and D. Marcus. Due process and mass adjudication: crisis and reform. Stanford Law Review, 72:1, 2020.
  11. 11.C. Anderson, C. Redcross, E. Valentine, and L. Miratrix. Evaluation of pretrial justice system reforms that use the public safety assessment: Effects of new jersey’s criminal justice reform. New York City: MDRC, 2019.
  12. 12.J. Angwin, J. Larson, S. Mattu, and L. Kirchner. Machine bias. ProPublica, May 2016. URL https://www.propublica.org/article/machine-bias-risk-assessments-in-criminal-sentencing.
  13. 13.APPR. Homepage — Advancing Pretrial Policy & Research (APPR) — advancingpretrial.org. https://advancingpretrial.org/improving-pretrial-justice/implement-pretrial-improvements/use-the-least-restrictive-conditions/, 2024. [Accessed 30-03-2025].
  14. 14.N. Arnosti and P. Shi. Design of lotteries and wait-lists for affordable housing allocation. Management Science, 66(6):2291–2307, 2020.
  15. 15.S. Athey and S. Wager. Policy learning with observational data. Econometrica, 89(1):133–161, 2021.
  16. 16.S. Athey, N. Keleher, and J. Spiess. Machine learning who to nudge: causal vs predictive targeting in a field experiment on student financial aid renewal. Journal of Econometrics, page 105945, 2025.
  17. 17.S. Baird, J. A. Bohren, C. McIntosh, and B. Özler. Optimal design of experiments in the presence of interference. Review of Economics and Statistics, 100(5):844–860, 2018.
  18. 18.J. Banasik and J. Crook. Reject inference, augmentation, and sample selection. European Journal of Operational Research, 183(3):1582–1594, 2007.
  19. 19.G. Bansal, B. Nushi, E. Kamar, D. S. Weld, W. S. Lasecki, and E. Horvitz. Updates in human-ai teams: Understanding and addressing the performance/compatibility tradeoff. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 2429–2437, 2019.
  20. 20.G. Bansal, T. Wu, J. Zhou, R. Fok, B. Nushi, E. Kamar, M. T. Ribeiro, and D. Weld. Does the whole exceed its parts? the effect of ai explanations on complementary team performance. In Proceedings of the 2021 CHI conference on human factors in computing systems, pages 1–16, 2021.
  21. 21.M. Bao, A. Zhou, S. Zottola, B. Brubach, S. Desmarais, A. Horowitz, K. Lum, and S. Venkatasubramanian. It’s compaslicated: The messy relationship between rai datasets and algorithmic fairness benchmarks. arXiv preprint arXiv:2106.05498, 2021.
  22. 22.C. Barabas, M. Virza, K. Dinakar, J. Ito, and J. Zittrain. Interventions over predictions: Reframing the ethical debate for actuarial risk assessment. In Conference on fairness, accountability and transparency, pages 62–76. PMLR, 2018.
  23. 23.S. Barocas, M. Hardt, and A. Narayanan. Fairness and machine learning: Limitations and opportunities. MIT press, 2023.
  24. 24.S. Basu. Using measures of race to make clinical predictions. Proceedings of the National Academy of Sciences, 120(25):e2303370120, 2023. doi: 10.1073/pnas.2303370120.
  25. 25.M. S. Bauer, L. Damschroder, H. Hagedorn, J. Smith, and A. M. Kilbourne. An introduction to implementation science for the non-specialist. BMC psychology, 3:1–12, 2015.
  26. 26.S. Behncke, M. Frölich, and M. Lechner. Targeting labour market programmes—results from a randomized experiment. Swiss Journal of Economics and Statistics, 145:221–268, 2009.
  27. 27.E. Ben-Michael, D. J. Greiner, M. Huang, K. Imai, Z. Jiang, and S. Shin. Does AI help humans make better decisions? a methodological framework for experimental evaluation. arXiv preprint arXiv:2403.12108, 2024.
  28. 28.E. Ben-Michael, D. J. Greiner, K. Imai, and Z. Jiang. Safe policy learning through extrapolation: Application to pre-trial risk assessment. Journal of the American Statistical Association, Forthcoming, 2025.
  29. 29.M. C. Berger, D. Black, and J. A. Smith. Evaluating profiling as a means of allocating government services. In M. Lechner and F. Pfeiffer, editors, Econometric Evaluation of Labour Market Policies, pages 59–84, Heidelberg, 2001. Physica-Verlag HD. ISBN 978-3-642-57615-7.
  30. 30.P. Bergman, R. Chetty, S. DeLuca, N. Hendren, L. F. Katz, and C. Palmer. Creating moves to opportunity: Experimental evidence on barriers to neighborhood choice. American Economic Review, 114(5):1281–1337, 2024.
  31. 31.C. Beyond. Why am i always being researched? https://wp.chicagobeyond.org/wp-content/uploads/2023/09/ChicagoBeyond_Why-Am-I.pdf, 2023.
  32. 32.U. Bhatt, J. Antorán, Y. Zhang, Q. V. Liao, P. Sattigeri, R. Fogliato, G. Melançon, R. Krishnan, J. Stanley, O. Tickoo, et al. Uncertainty as a form of transparency: Measuring, communicating, and using uncertainty. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pages 401–413, 2021.
  33. 33.A. D. Biderman and A. J. Reiss Jr. On exploring the” dark figure” of crime. The Annals of the American Academy of Political and Social Science, 374(1):1–15, 1967.
  34. 34.D. A. Black, J. A. Smith, M. C. Berger, and B. J. Noel. Is the threat of reemployment services more effective than the services themselves? evidence from random assignment in the ui system. The American Economic Review, 2003.
  35. 35.W. Boag, A. Hasan, J. Y. Kim, M. Revoir, M. Nichols, W. Ratliff, M. Gao, S. Zilberstein, Z. Samad, Z. Hoodbhoy, et al. The algorithm journey map: a tangible approach to implementing ai solutions in healthcare. NPJ Digital Medicine, 7(1):87, 2024.
  36. 36.A. Boussina, S. P. Shashikumar, A. Malhotra, R. L. Owens, R. El-Kareh, C. A. Longhurst, K. Quintero, A. Donahue, T. C. Chan, S. Nemati, et al. Impact of a deep learning sepsis prediction model on quality of care and survival. NPJ digital medicine, 7(1):14, 2024.
  37. 37.J. Boutilier, J. O. Jonasson, H. Li, and E. Yoeli. Randomized controlled trials of service interventions: The impact of capacity constraints. arXiv preprint arXiv:2407.21322, 2024.
  38. 38.A. Boyarsky, H. Namkoong, and J. Pouget-Abadie. Modeling interference using experiment roll-out. arXiv preprint arXiv:2305.10728, 2023.
  39. 39.C. Brown, M. Ravallion, and D. van de Walle. A poor means test? econometric targeting in africa. Policy Research Working Paper 7915: World Bank Group, Development Research Group, Human Development and Public Services Team, 2016.
  40. 40.Z. Buçinca, P. Lin, K. Z. Gajos, and E. L. Glassman. Proxy tasks and subjective measures can be misleading in evaluating explainable ai systems. In Proceedings of the 25th international conference on intelligent user interfaces, pages 454–464, 2020.
  41. 41.Z. Buçinca, M. B. Malaya, and K. Z. Gajos. To trust or to think: cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making. Proceedings of the ACM on Human-computer Interaction, 5(CSCW1):1–21, 2021.
  42. 42.G. Burke. Child welfare algorithm faces Justice Department scrutiny — apnews.com. https://apnews.com/article/justice-scrutinizes-pittsburgh-child-welfare-ai-tool-4f61f45bfc3245fd2556e886c2da988b. [Accessed 10-11-2024].
  43. 43.K. S. Button, J. P. Ioannidis, C. Mokrysz, B. A. Nosek, J. Flint, E. S. Robinson, and M. R. Munafò. Power failure: why small sample size undermines the reliability of neuroscience. Nature reviews neuroscience, 14(5):365–376, 2013.
  44. 44.R. Caruana, Y. Lou, J. Gehrke, P. Koch, M. Sturm, and N. Elhadad. Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission. In Proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining, pages 1721–1730, 2015.
  45. 45.L. W. Chang, E. L. Kirgios, S. Mullainathan, and K. L. Milkman. Does counting change what counts? Quantification fixation biases decision-making. Proceedings of the National Academy of Sciences, 121(46):e2400215121, Nov. 2024. doi: 10.1073/pnas.2400215121. URL https://www.pnas.org/doi/10.1073/pnas.2400215121. Publisher: Proceedings of the National Academy of Sciences.
  46. 46.R. N. Charette. Michigan’s midas unemployment system: Algorithm alchemy created lead, not gold. IEEE Spectrum, 24(3), 2018.
  47. 47.L. Cheng, C. Drayton, A. Chouldechova, and R. Vaithianathan. Algorithm-assisted decision making and racial disparities in housing: A study of the allegheny housing assessment tool. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 281–292, 2024.
  48. 48.J. Chien, M. Roberts, and B. Ustun. Algorithmic censoring in dynamic learning systems. In Proceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, pages 1–20, 2023.
  49. 49.A. Chouldechova, D. Benavides-Prado, O. Fialko, and R. Vaithianathan. A case study of algorithm-assisted decision making in child maltreatment hotline screening decisions. In Conference on fairness, accountability and transparency, pages 134–148. PMLR, 2018.
  50. 50.J. Chu. Cameras of Merit or Engines of Inequality? College Ranking Systems and the Enrollment of Disadvantaged Students. American Journal of Sociology, 126(6):1307–1346, May 2021. ISSN 0002-9602. doi: 10.1086/714916. URL https://www-journals-uchicago-edu.proxy.library.nyu.edu/doi/10.1086/714916. Publisher: The University of Chicago Press.
  51. 51.M. Coots, S. Basu, and Z. Obermeyer. Racial bias in clinical and population health algorithms. Annual Review of Public Health, 46:1–19, 2023. doi: 10.1146/annurev-publhealth-071823-112058.
  52. 52.M. Coots, S. Saghafian, D. M. Kent, and S. Goel. A framework for considering the value of race and ethnicity in estimating disease risk. Annals of Internal Medicine, 178(1):98–107, 2025.
  53. 53.S. Corbett-Davies and S. Goel. The Measure and Mismeasure of Fairness: A Critical Review of Fair Machine Learning. CoRR, abs/1808.00023, 2018.
  54. 54.S. Corbett-Davies, E. Pierson, A. Feller, S. Goel, and A. Huq. Algorithmic decision making and the cost of fairness. In Proceedings of the 23rd acm sigkdd international conference on knowledge discovery and data mining, pages 797–806, 2017.
  55. 55.C. Cortes, G. DeSalvo, and M. Mohri. Learning with rejection. In Algorithmic Learning Theory: 27th International Conference, ALT 2016, Bari, Italy, October 19-21, 2016, Proceedings 27, pages 67–82. Springer, 2016.
  56. 56.N. Corvelo Benz and M. Rodriguez. Human-aligned calibration for ai-assisted decision making. Advances in Neural Information Processing Systems, 36, 2024.
  57. 57.A. Coston, A. Mishler, E. H. Kennedy, and A. Chouldechova. Counterfactual risk assessments, evaluation, and fairness. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pages 582–593, 2020.
  58. 58.A. Coston, A. Rambachan, and A. Chouldechova. Characterizing fairness over the set of good models under selective labels. In International Conference on Machine Learning, pages 2144–2155. PMLR, 2021.
  59. 59.A. Coston, A. Kawakami, H. Zhu, K. Holstein, and H. Heidari. A validity perspective on evaluating the justified use of data-driven decision-making algorithms. In 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), pages 690–704. IEEE, 2023.
  60. 60.B. Cowgill. The impact of algorithms on judicial discretion: Evidence from regression discontinuities. Unpublished Manuscript, Columbia Business School, 2018.
  61. 61.B. Cowgill and M. T. Stevenson. Algorithmic social engineering. In AEA Papers and Proceedings, volume 110, pages 96–100. American Economic Association 2014 Broadway, Suite 305, Nashville, TN 37203, 2020.
  62. 62.K. Creel and D. Hellman. The algorithmic leviathan: Arbitrariness, fairness, and opportunity in algorithmic decision-making systems. Canadian Journal of Philosophy, 52(1):26–43, 2022.
  63. 63.R. K. Crump, V. J. Hotz, G. W. Imbens, and O. A. Mitnik. Nonparametric tests for treatment effect heterogeneity. The Review of Economics and Statistics, 90(3):389–405, 2008.
  64. 64.M. De-Arteaga, A. Dubrawski, and A. Chouldechova. Learning under selective labels in the presence of expert consistency. Workshop on Fairness, Accountability, and Transparency in Machine Learning (FAT/ML), 2018.
  65. 65.J. L. Doleac. Study after study shows ex-prisoners would be better off without intense supervision. The Brookings Institution, 2018.
  66. 66.D. Donoho. Data science at the singularity. Harvard Data Science Review, 6(1), 2024.
  67. 67.D. Dranove, D. Kessler, M. McClellan, and M. Satterthwaite. Is More Information Better? The Effects of “Report Cards” on Health Care Providers. Journal of Political Economy, 111(3):555–588, June 2003. ISSN 0022-3808. doi: 10.1086/374180. URL https://www.journals.uchicago.edu/doi/abs/10.1086/374180. Publisher: The University of Chicago Press.
  68. 68.M. Dudík, J. Langford, and L. Li. Doubly robust policy evaluation and learning. arXiv preprint arXiv:1103.4601, 2011.
  69. 69.C. Dwork and C. Ilvento. Individual fairness under composition. Proceedings of fairness, accountability, transparency in machine learning, 2018.
  70. 70.L. Eckhouse, K. Lum, C. Conti-Cook, and J. Ciccolini. Layers of bias: A unified approach for understanding problems with risk assessment. Criminal Justice and Behavior, 46(2):185–209, 2018. doi: 10.1177/0093854818811379.
  71. 71.A. Ehrhardt, C. Biernacki, V. Vandewalle, P. Heinrich, and S. Beben. Reject inference methods in credit scoring. Journal of Applied Statistics, 48(13-15):2734–2754, 2021.
  72. 72.H. Elzayn, E. Smith, T. Hertz, C. Guage, A. Ramesh, R. Fisher, D. E. Ho, and J. Goldin. Measuring and mitigating racial disparities in tax audits*. The Quarterly Journal of Economics, 140(1):113–163, 09 2024. ISSN 0033-5533. doi: 10.1093/qje/qjae027. URL https://doi.org/10.1093/qje/qjae027.
  73. 73.A. Engler. Enrollment algorithms are contributing to the crises of higher education. Technical report, The Brookings Institution, Sept. 2021.
  74. 74.W. N. Espeland and M. L. Stevens. A Sociology of Quantification. European Journal of Sociology / Archives Européennes de Sociologie, 49(3):401–436, Dec. 2008. ISSN 1474-0583, 0003-9756. doi: 10.1017/S0003975609000150. URL https://www.cambridge.org/core/journals/european-journal-of-sociology-archives-europeennes-de-sociologie/article/sociology-of-quantification/53D563E0E4A75A05E877B27C06E957F9.
  75. 75.T. Feathers. Major Universities Are Using Race as a High-Impact Predictor of Student Success. The Markup, March 2021. URL https://themarkup.org/machine-learning/2021/03/02/major-universities-are-using-race-as-a-high-impact-predictor-of-student-success. Accessed: 2025-02-11.
  76. 76.T. Feathers. Takeaways from our investigation into wisconsin’s racially inequitable dropout algorithm. The Markup, April 27 2023. URL https://themarkup.org/the-breakdown/2023/04/27/takeaways-from-our-investigation-into-wisconsins-racially-inequitable-dropout-algorithm.
  77. 77.T. Felin, M. Sako, and J. Hullman. Artificial intelligence and actor-specific decisions. Available at SSRN 5279401, 2025.
  78. 78.U. Fischer-Abaigar, C. Kern, and J. C. Perdomo. The value of prediction in identifying the worst-off. International Conference on Machine Learning, 2025.
  79. 79.R. Fogliato, A. Chouldechova, and M. G’Sell. Fairness evaluation in presence of biased noisy labels. In International conference on artificial intelligence and statistics, pages 2325–2336. PMLR, 2020.
  80. 80.R. J. Gallo, L. Shieh, M. Smith, B. J. Marafino, P. Geldsetzer, S. M. Asch, K. Shum, S. Lin, J. Westphal, G. Hong, et al. Effectiveness of an artificial intelligence–enabled intervention for detecting clinical deterioration. JAMA internal medicine, 184(5):557–562, 2024.
  81. 81.M. Gerchick, T. Jegede, T. Shah, A. Gutierrez, S. Beiers, N. Shemtov, K. Xu, A. Samant, and A. Horowitz. The devil is in the details: Interrogating values embedded in the allegheny family screening tool. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, pages 1292–1310, 2023.
  82. 82.K. Goddard, A. Roudsari, and J. C. Wyatt. Automation bias: a systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association, 19(1):121–127, 2012.
  83. 83.C. A. Goodhart. Problems of monetary management: the UK experience. Springer, 1984.
  84. 84.T. Gowan. Hobos, hustlers, and backsliders: Homeless in San Francisco. U of Minnesota Press, 2010.
  85. 85.B. Green. The false promise of risk assessments: Epistemic reform and the limits of fairness. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (FAT* 2020), pages 594–606. Association for Computing Machinery, 2020. doi: 10.1145/3351095.3372869. URL https://dl.acm.org/doi/10.1145/3351095.3372869.
  86. 86.B. Green and Y. Chen. Algorithmic risk assessments can alter human decision-making processes in high-stakes government contexts. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW2):1–33, 2021.
  87. 87.N. Grgić-Hlača, G. Lima, A. Weller, and E. M. Redmiles. Dimensions of Diversity in Human Perceptions of Algorithmic Fairness. In Proceedings of the 2nd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, EAAMO ’22, pages 1–12, New York, NY, USA, Oct. 2022. Association for Computing Machinery. ISBN 978-1-4503-9477-2. doi: 10.1145/3551624.3555306. URL https://dl.acm.org/doi/10.1145/3551624.3555306.
  88. 88.W. M. Grove, D. H. Zald, B. S. Lebow, B. E. Snitz, and C. Nelson. Clinical versus mechanical prediction: a meta-analysis. Psychological assessment, 12(1):19, 2000.
  89. 89.L. Guerdan, A. Coston, K. Holstein, and Z. S. Wu. Counterfactual prediction under outcome measurement error. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, pages 1584–1598, 2023.
  90. 90.L. Guerdan, A. L. Coston, K. Holstein, and S. Wu. Predictive performance comparison of decision policies under confounding. In International Conference on Machine Learning, pages 16673–16705. PMLR, 2024.
  91. 91.L. Guerdan, D. Saxena, S. Chancellor, Z. S. Wu, and K. Holstein. Measurement as bricolage: How data scientists construct target variables for predictive modeling tasks. Proceedings of the ACM on Human-Computer Interaction, (CSCW), 2025. Forthcoming.
  92. 92.Z. Guo, Y. Wu, J. D. Hartline, and J. Hullman. A decision theoretic framework for measuring ai reliance. In The 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 221–236, 2024.
  93. 93.Z. Guo, Y. Wu, J. Hartline, and J. Hullman. The value of information in human-ai decision-making. arXiv preprint arXiv:2502.06152, 2025.
  94. 94.M. Hardt and C. Mendler-Dunner. Performative prediction: Past and future. Statistical Science, 2025.
  95. 95.M. Hardt, N. Megiddo, C. Papadimitriou, and M. Wootters. Strategic classification. In Proceedings of the 2016 ACM conference on innovations in theoretical computer science, pages 111–122, 2016.
  96. 96.M. Hardt, M. Jagadeesan, and C. Mendler-Dünner. Performative power. Advances in Neural Information Processing Systems, 2022.
  97. 97.J. Haushofer, P. Niehaus, C. Paramo, E. Miguel, and M. W. Walker. Targeting impact versus deprivation. Technical report, National Bureau of Economic Research, 2022.
  98. 98.M. A. Hernán and J. M. Robins. Causal inference.
  99. 99.D. E. Ho, C. Handan-Nader, D. Ames, and D. Marcus. Quality review of mass adjudication: A randomized natural experiment at the board of veterans appeals, 2003–16. The Journal of Law, Economics, and Organization, 35(2):239–288, 03 2019.
  100. 100.J. Holl, G. Kernbeiß, and M. Wagner-Pinter. Das ams-arbeitsmarktchancen-modell. Arbeitsmarktservice Österreich, Wien, 2018.
  101. 101.S. P. Horbach, J. K. Tijdink, and L. M. Bouter. Partial lottery can make grant allocation more fair, more efficient, and more diverse. Science and Public Policy, 49(4):580–582, 2022.
  102. 102.L. Hu. What is “race” in algorithmic discrimination on the basis of race? Journal of Moral Philosophy, 1(aop):1–26, 2023.
  103. 103.L. Hu and Y. Chen. A short-term intervention for long-term fairness in the labor market. In Proceedings of the 2018 World Wide Web Conference, pages 1389–1398, 2018.
  104. 104.L. Hu and Y. Chen. Fair classification and social welfare. In Proceedings of the 2020 conference on fairness, accountability, and transparency, pages 535–545, 2020.
  105. 105.L. Hu and I. Kohler-Hausmann. What’s sex got to do with machine learning? In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pages 513–513, 2020.
  106. 106.L. Hu, N. Immorlica, and J. W. Vaughan. The disparate effects of strategic manipulation. In Proceedings of the Conference on Fairness, Accountability, and Transparency, FAT* ’19, pages 259–268, New York, NY, USA, 2019. ACM. ISBN 978-1-4503-6125-5. doi: 10.1145/3287560.3287597.
  107. 107.J. Hullman, Z. Guo, and B. Ustun. Explanations are a means to an end. arXiv preprint, 2025a.
  108. 108.J. Hullman, A. Kale, and J. Hartline. Underspecified human decision experiments considered harmful. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI ’25, New York, NY, USA, 2025b. Association for Computing Machinery. ISBN 9798400713941. doi: 10.1145/3706598.3714063. URL https://doi.org/10.1145/3706598.3714063.
  109. 109.J. Hullman, Y. Wu, D. Xie, Z. Guo, and A. Gelman. Conformal prediction and human decision making. arXiv preprint arXiv:2503.11709, 2025c.
  110. 110.K. Imai and Z. Jiang. Principal fairness for human and algorithmic decision-making. Statistical Science, 38(2):317–328, 2023.
  111. 111.K. Imai, Z. Jiang, J. Greiner, R. Halen, and S. Shin. Experimental evaluation of algorithm-assisted human decision-making: Application to pretrial public safety assessment (with discussion). Journal of the Royal Statistical Society, Series A (Statistics in Society), 186(2):167–189, April 2023.
  112. 112.G. W. Imbens and D. B. Rubin. Causal inference in statistics, social, and biomedical sciences. Cambridge University Press, 2015.
  113. 113.A. Z. Jacobs and H. Wallach. Measurement and fairness. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 375–385, 2021.
  114. 114.R. A. Johnson and S. Zhang. What is the bureaucratic counterfactual? categorical versus algorithmic prioritization in us social policy. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 1671–1682, 2022.
  115. 115.R. A. Johnson and S. Zhang. Predictive algorithms and perceptions of fairness: Parent attitudes toward algorithmic resource allocation in k-12 education. Sociological Science, 12:322–356, May 2025. doi: 10.15195/v12.a15. URL https://sociologicalscience.com/articles-v12-15-322/.
  116. 116.I. D. Jong. The Time Seems Right: Let’s Begin the End of the VI-SPDAT — OrgCode Consulting — orgcode.com. https://www.orgcode.com/blog/the-time-seems-right-lets-begin-the-end-of-the-vi-spdat. [Accessed 01-05-2025].
  117. 117.D. Kahneman. Maps of bounded rationality: Psychology for behavioral economics. American economic review, 93(5):1449–1475, 2003.
  118. 118.N. Kallus and A. Zhou. Residual unfairness in fair machine learning from prejudiced data. In J. Dy and A. Krause, editors, Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 2439–2448, Stockholmsmässan, Stockholm Sweden, 10–15 Jul 2018. PMLR.
  119. 119.N. Kallus and A. Zhou. Assessing disparate impact of personalized interventions: identifiability and bounds. Advances in neural information processing systems, 32, 2019.
  120. 120.N. Kallus and A. Zhou. Minimax-optimal policy learning under unobserved confounding. Management Science, 67(5):2870–2890, 2021.
  121. 121.M. Katell, M. Young, B. Herman, D. Dailey, A. Tam, V. Guetler, C. Binz, D. Raz, and P. Krafft. An algorithmic equity toolkit for technology audits by community advocates and activists. arXiv preprint arXiv:1912.02943, 2019.
  122. 122.A. Kawakami, A. Coston, H. Zhu, H. Heidari, and K. Holstein. The situate ai guidebook: Co-designing a toolkit to support multi-stakeholder, early-stage deliberations around public sector ai proposals. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 1–22, 2024.
  123. 123.B. Kelly and D. F. Perkins. Handbook of implementation science for psychology in education. Cambridge University Press, 2012.
  124. 124.E. H. Kennedy. Semiparametric theory and empirical processes in causal inference. Statistical causal inferences and their applications in public health research, pages 141–167, 2016.
  125. 125.R. H. Keogh and N. Van Geloven. Prediction under interventions: evaluation of counterfactual performance using longitudinal observational data. Epidemiology, 35(3):329–339, 2024.
  126. 126.N. Kilbertus, M. Rojas Carulla, G. Parascandolo, M. Hardt, D. Janzing, and B. Schölkopf. Avoiding discrimination through causal reasoning. Advances in neural information processing systems, 30, 2017a.
  127. 127.N. Kilbertus, M. Rojas-Carulla, G. Parascandolo, M. Hardt, D. Janzing, and B. Schölkopf. Avoiding discrimination through causal reasoning. 2017b.
  128. 128.A. Kim, M. Yang, and J. Zhang. When algorithms err: Differential impact of early vs. late errors on users’ reliance on algorithms. ACM Transactions on Computer-Human Interaction, 30(1):1–36, 2023.
  129. 129.M. P. Kim and J. C. Perdomo. Making decisions under outcome performativity. Innovations in Theoretical Computer Science, 2023.
  130. 130.M. Kitzmiller and M. Gewirtz. Validation of thenew york city criminaljustice agency pretrialrelease assessment. https://www.nycja.org/assets/downloads/Release-Assessment-Validation-1.0-20250512-2.pdf, 2025. [Accessed 19-10-2025].
  131. 131.G. A. Klein. Sources of power: How people make decisions. MIT press, 2017.
  132. 132.J. Kleinberg and M. Raghavan. How Do Classifiers Induce Agents to Invest Effort Strategically? In Proceedings of the 2019 ACM Conference on Economics and Computation, EC ’19, pages 825–844, New York, NY, USA, 2019. ACM. ISBN 978-1-4503-6792-9.
  133. 133.J. Kleinberg and M. Raghavan. Algorithmic monoculture and social welfare. Proceedings of the National Academy of Sciences, 118(22):e2018340118, 2021.
  134. 134.J. Kleinberg, J. Ludwig, S. Mullainathan, and Z. Obermeyer. Prediction policy problems. American Economic Review, 105(5):491–95, 2015.
  135. 135.J. Kleinberg, J. Ludwig, and S. Mullainathan. A guide to solving social problems with machine learning. Harvard Business Review, December 2016.
  136. 136.J. Kleinberg, H. Lakkaraju, J. Ludwig, J. Leskovec, and S. Mullanaithan. Human decisions and machine predictions. National Bureau of Economic Research, 2017.
  137. 137.J. Kleinberg, H. Lakkaraju, J. Leskovec, J. Ludwig, and S. Mullainathan. Human decisions and machine predictions. The quarterly journal of economics, 133(1):237–293, 2018.
  138. 138.J. E. Knowles. Of needles and haystacks: Building an accurate statewide dropout early warning system in wisconsin. Journal of Educational Data Mining, 7(3):18–67, 2015. doi: 10.5281/zenodo.3554725. URL https://jedm.educationaldatamining.org/index.php/JEDM/article/view/JEDM082.
  139. 139.J. L. Koepke and D. G. Robinson. Danger ahead: Risk assessment and the future of bail reform. Wash. L. Rev., 93:1725, 2018.
  140. 140.A. Kofman. Digital jail: How electronic monitoring drives defendants into debt. The New York Times Magazine, 3, 2019.
  141. 141.S. Kross and P. Guo. Orienting, framing, bridging, magic, and counseling: How data scientists navigate the outer loop of client collaborations in industry and academia. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW2):1–28, 2021.
  142. 142.A. Kube and S. Das. Allocating interventions based on predicted outcomes: A case study on homelessness services. Proceedings of the AAAI Conference on Artificial Intelligence, 2019.
  143. 143.A. Kube, S. Das, and P. Fowler. Community- and data-driven homelessness prevention and service delivery: optimizing for equity. Journal of the American Medical Informatics Association, 30(6):1032–1041, 2023a. doi: 10.1093/jamia/ocad052.
  144. 144.A. R. Kube, S. Das, and P. J. Fowler. Fair and Efficient Allocation of Scarce Resources Based on Predicted Outcomes: Implications for Homeless Service Delivery. J. Artif. Int. Res., 76, 2023b. doi: 10.1613/jair.1.12847.
  145. 145.M. J. Kusner, J. Loftus, C. Russell, and R. Silva. Counterfactual fairness. Advances in neural information processing systems, 30, 2017.
  146. 146.H. Lakkaraju, J. Kleinberg, J. Leskovec, J. Ludwig, and S. Mullainathan. The selective labels problem: Evaluating algorithmic predictions in the presence of unobservables. In Proceedings of the 23rd ACM SIGKDD international conference on knowledge discovery and data mining, pages 275–284, 2017.
  147. 147.B. Laufer and H. Nissenbaum. Algorithmic displacement of social trust. Knight First Amendment Institute at Columbia University, 2023.
  148. 148.B. Laufer, T. K. Gilbert, and H. Nissenbaum. Optimization’s Neglected Normative Commitments. In 2023 ACM Conference on Fairness, Accountability, and Transparency, pages 50–63, June 2023a. doi: 10.1145/3593013.3593976. URL http://arxiv.org/abs/2305.17465. arXiv:2305.17465 [cs].
  149. 149.B. Laufer, J. Kleinberg, K. Levy, and H. Nissenbaum. Strategic evaluation. In Proceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization, pages 1–12, 2023b.
  150. 150.J. J. Lee, R. Srinivasan, C. S. Ong, D. Alejo, S. Schena, I. Shpitser, M. Sussman, G. J. Whitman, and D. Malinsky. Causal determinants of postoperative length of stay in cardiac surgery using causal graphical learning. The Journal of Thoracic and Cardiovascular Surgery, 166(5):e446–e462, 2023.
  151. 151.J. Levy, M. van der Laan, A. Hubbard, and R. Pirracchio. A fundamental measure of treatment effect heterogeneity. Journal of Causal Inference, 9(1):83–108, 2021a.
  152. 152.K. Levy, K. E. Chasalow, and S. Riley. Algorithms and decision-making in the public sector. Annual Review of Law and Social Science, 17(1):309–334, 2021b.
  153. 153.L. T. Liu, S. Dean, E. Rolf, M. Simchowitz, and M. Hardt. Delayed impact of fair machine learning. In J. Dy and A. Krause, editors, Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 3150–3158, Stockholm, Sweden, 10–15 Jul 2018. PMLR.
  154. 154.L. T. Liu, A. Wilson, N. Haghtalab, A. T. Kalai, C. Borgs, and J. Chayes. The disparate equilibria of algorithmic decision making when individuals invest rationally. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, pages 381–391, 2020.
  155. 155.L. T. Liu, N. Garg, and C. Borgs. Strategic ranking. In International Conference on Artificial Intelligence and Statistics, pages 2489–2518. PMLR, 2022.
  156. 156.L. T. Liu, S. Barocas, J. Kleinberg, and K. Levy. On the actionability of outcome prediction. arXiv preprint arXiv:2309.04470, 2023a.
  157. 157.L. T. Liu, S. Wang, T. Britton, and R. Abebe. Reimagining the machine learning life cycle to improve educational outcomes of students. Proceedings of the National Academy of Sciences, 120(9):e2204781120, 2023b.
  158. 158.J. R. Loftus, L. E. Bynum, and S. Hansen. Causal dependence plots. arXiv preprint arXiv:2303.04209, 2023.
  159. 159.J. M. Logg, J. A. Minson, and D. A. Moore. Algorithm appreciation: People prefer algorithmic to human judgment. Organizational Behavior and Human Decision Processes, 151:90–103, 2019.
  160. 160.Z. Lu and M. Yin. Human reliance on machine learning models when performance feedback is limited: Heuristics and risks. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, pages 1–16, 2021.
  161. 161.J. Ludwig, S. Mullainathan, and A. Rambachan. The unreasonable effectiveness of algorithms. In AEA Papers and Proceedings, volume 114, pages 623–627. American Economic Association 2014 Broadway, Suite 305, Nashville, TN 37203, 2024.
  162. 162.I. Lundberg, R. Brown-Weinstock, S. Clampet-Lundquist, S. Pachman, T. J. Nelson, V. Yang, K. Edin, and M. J. Salganik. The origins of unpredictability in life outcome prediction tasks. Proceedings of the National Academy of Sciences, 121(24):e2322973121, 2024.
  163. 163.G. Maheshwari, A. Bellet, P. Denis, and M. Keller. Fair without leveling down: A new intersectional fairness definition. arXiv preprint arXiv:2305.12495, 2023.
  164. 164.M. Makar, B. Packer, D. Moldovan, D. Blalock, Y. Halpern, and A. D’Amour. Causally motivated shortcut removal using auxiliary labels. In International Conference on Artificial Intelligence and Statistics, pages 739–766. PMLR, 2022.
  165. 165.C. F. Manski. Statistical treatment rules for heterogeneous populations. Econometrica, 72(4):1221–1246, 2004.
  166. 166.C. F. Manski. Probabilistic prediction for binary treatment choice: With focus on personalized medicine. Journal of Econometrics, 234(2):647–663, 2023. doi: 10.1016/j.jeconom.2022.07.009.
  167. 167.K. Martin and A. Waldman. Are Algorithmic Decisions Legitimate? The Effect of Process and Outcomes on Perceptions of Legitimacy of AI Decisions. Journal of Business Ethics, 183(3):653–670, Mar. 2023. ISSN 0167-4544, 1573-0697. doi: 10.1007/s10551-021-05032-7. URL https://link.springer.com/10.1007/s10551-021-05032-7.
  168. 168.S. Mathison. Encyclopedia of evaluation. Sage publications, 2004.
  169. 169.K. McConvey, S. Guha, and A. Kuzminykh. A human-centered review of algorithms in decision-making in higher education. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–15, 2023.
  170. 170.B. McLaughlin and J. Spiess. Algorithmic assistance with recommendation-dependent preferences. arXiv preprint arXiv:2208.07626, 2022.
  171. 171.B. McLaughlin and J. Spiess. Designing algorithmic recommendations to achieve human-ai complementarity. arXiv preprint arXiv:2405.01484, 2024.
  172. 172.P. E. Meehl. Clinical versus statistical prediction: A theoretical analysis and a review of the evidence. 1954.
  173. 173.C. Mendler-Dünner, F. Ding, and Y. Wang. Anticipating performativity by predicting from predictions. Advances in Neural Information Processing Systems, 35:31171–31185, 2022.
  174. 174.M. Meyer, A. Horowitz, E. Marshall, and K. Lum. Flipping the script on criminal justice risk assessment: An actuarial model for assessing the risk the federal sentencing system poses to defendants. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 366–378, 2022.
  175. 175.J. M. Mikhaeil, A. Gelman, and P. Greengard. Hierarchical bayesian models to mitigate systematic disparities in prediction with proxy outcomes. Journal of the Royal Statistical Society Series A: Statistics in Society, page qnae142, 2024.
  176. 176.J. Miller, S. Milli, and M. Hardt. Strategic classification is causal modeling in disguise. In International Conference on Machine Learning, pages 6917–6926. PMLR, 2020.
  177. 177.J. P. Miller, J. C. Perdomo, and T. Zrnic. Outside the echo chamber: Optimizing the performative risk. In International Conference on Machine Learning, pages 7710–7720. PMLR, 2021.
  178. 178.S. Milli, J. Miller, A. D. Dragan, and M. Hardt. The social cost of strategic classification. In Proceedings of the Conference on Fairness, Accountability, and Transparency, FAT* ’19, pages 230–239, New York, NY, USA, 2019. ACM. ISBN 978-1-4503-6125-5. doi: 10.1145/3287560.3287576.
  179. 179.M. Mitchell, S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru. Model cards for model reporting. In Proceedings of the conference on fairness, accountability, and transparency, pages 220–229, 2019.
  180. 180.B. Mittelstadt, S. Wachter, and C. Russell. The unfairness of fair machine learning: Leveling down and strict egalitarianism by default. Mich. Tech. L. Rev., 30:1, 2023.
  181. 181.T. Moir. Why is implementation science important for intervention design and evaluation within educational settings? In Frontiers in Education, volume 3, page 61. Frontiers Media SA, 2018.
  182. 182.R. C. Moore. Navigating employment discrimination in ai and automated systems: A new civil rights frontier. Testimony before the U.S. Equal Employment Opportunity Commission, Jan. 2023. URL https://www.aclu.org/wp-content/uploads/2023/10/ACLU-testimony-to-EEOC-on-Employment-AI-for-Jan-31-2023-hearing2-1.docx.pdf?utm_source=chatgpt.com. Accessed: 2025-04-18.
  183. 183.H. Mozannar and D. Sontag. Consistent estimators for learning to defer to an expert. In International conference on machine learning, pages 7076–7087. PMLR, 2020.
  184. 184.H. Mozannar, A. Satyanarayan, and D. Sontag. Teaching humans when to defer to a classifier via exemplars. In Proceedings of the aaai conference on artificial intelligence, volume 36, pages 5323–5331, 2022.
  185. 185.H. Mozannar, H. Lang, D. Wei, P. Sattigeri, S. Das, and D. Sontag. Who should predict? exact algorithms for learning to defer to humans. In International conference on artificial intelligence and statistics, pages 10520–10545. PMLR, 2023.
  186. 186.S. Mullainathan. Biased algorithms are easier to fix than biased people. The New York Times, 2019.
  187. 187.S. Mullainathan and Z. Obermeyer. On the inequity of predicting a while hoping for b. In AEA Papers and Proceedings, volume 111, pages 37–42. American Economic Association 2014 Broadway, Suite 305, Nashville, TN 37203, 2021.
  188. 188.J. Z. Muller. The Tyranny of Metrics. Princeton University Press, Princeton, Feb. 2018. ISBN 978-0-691-17495-2.
  189. 189.E. Munro, S. Wager, and K. Xu. Treatment effects in market equilibrium. arXiv preprint arXiv:2109.11647, 2021.
  190. 190.R. Nabi and D. Benkeser. Fair risk minimization under causal path-specific effect constraints. arXiv preprint arXiv:2408.01630, 2024.
  191. 191.R. Nabi and I. Shpitser. Fair inference on outcomes. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 32, 2018.
  192. 192.R. Nabi, D. Malinsky, and I. Shpitser. Learning optimal fair policies. In Proceedings of the International Conference on Machine Learning, pages 4674–4682. PMLR, 2019.
  193. 193.R. Nabi, D. Malinsky, and I. Shpitser. Optimal training of fair predictive models. In Proceedings of the Conference on Causal Learning and Reasoning, pages 594–617. PMLR, 2022.
  194. 194.R. Nabi, N. S. Hejazi, M. J. van der Laan, and D. Benkeser. Statistical learning for constrained functional parameters in infinite-dimensional models with applications in fair machine learning. arXiv preprint arXiv:2404.09847, 2024.
  195. 195.S. Nagaraj, Y. Liu, F. Calmon, and B. Ustun. Regretful decisions under label noise. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=7B9FCDoUzB.
  196. 196.R. Neil and M. Zanger-Tishler. Algorithmic bias in criminal risk assessment: The consequences of racial differences in arrest as a measure of crime. Annual Review of Criminology, 8, 2024.
  197. 197.T. Nguyen, S. Alam, C. Hu, C. Albiston, and N. Salehi. Definitions of fairness differ across socioeconomic groups & shape perceptions of algorithmic decisions. Proceedings of the ACM on Human-Computer Interaction, 8(CSCW2):1–31, 2024. doi: 10.1145/3687058.
  198. 198.G. Noarov, R. Ramalingam, A. Roth, and S. Xie. High-dimensional prediction for sequential decision making. arXiv preprint arXiv:2310.17651, 2023.
  199. 199.Z. Obermeyer, B. Powers, C. Vogeli, and S. Mullainathan. Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464):447–453, 2019.
  200. 200.V. Ojewale, R. Steed, B. Vecchione, A. Birhane, and I. D. Raji. Towards ai accountability infrastructure: Gaps and opportunities in ai audit tooling. arXiv preprint arXiv:2402.17861, 2024.
  201. 201.Optum. Impact pro for analyzing future health risk, 2022. URL https://www.optum.com/business/health-plans/data-analytics/predict-health-risk.html.
  202. 202.F. Pasquale. The black box society: The secret algorithms that control money and information. Harvard University Press, 2015.
  203. 203.S. Passi and S. Barocas. Problem formulation and fairness. In Proceedings of the conference on fairness, accountability, and transparency, pages 39–48, 2019.
  204. 204.S. Passi and S. J. Jackson. Trust in data science: Collaboration, translation, and accountability in corporate data science projects. Proceedings of the ACM on human-computer interaction, 2(CSCW):1–28, 2018.
  205. 205.S. Passi and M. Vorvoreanu. Overreliance on ai literature review. Microsoft Research, 2022.
  206. 206.J. Perdomo, T. Zrnic, C. Mendler-Dünner, and M. Hardt. Performative prediction. In International Conference on Machine Learning, pages 7599–7609. PMLR, 2020.
  207. 207.J. C. Perdomo. The relative value of prediction in algorithmic decision making. International Conference on Machine Learning, 2024.
  208. 208.J. C. Perdomo. Revisiting the predictability of performative, social events, 2025.
  209. 209.J. C. Perdomo, T. Britton, M. Hardt, and R. Abebe. Difficult lessons on social prediction from wisconsin public schools. arXiv preprint arXiv:2304.06205, 2023.
  210. 210.L. Petry, C. Hill, P. Vayanos, E. Rice, H.-T. Hsu, and M. Morton. Associations between the vulnerability index-service prioritization decision assistance tool and returns to homelessness among single adults in the united states. Cityscape, 23(2):293–324, 2021.
  211. 211.T. Phillips. Ethics of field experiments. Annual Review of Political Science, 24(1):277–300, 2021.
  212. 212.T. M. Porter. Trust in Numbers: The Pursuit of Objectivity in Science and Public Life. Princeton University Press, Princeton, 1995. ISBN 978-1-4008-2161-7. URL https://princetonup-degruyter-com.ezproxy.princeton.edu/view/title/511899.
  213. 213.Pretrial Justice Institute. Scan of pretrial practices. Technical report, Pretrial Justice Institute, 2019. URL https://www.pretrial.org/files/resources/scanofpretrialpractices.pdf. Accessed: 2025-04-18.
  214. 214.M. Raghavan and P. T. Kim. Limitations of the” four-fifths rule” and statistical parity tests for measuring fairness. Geo. L. Tech. Rev., 8:93, 2024.
  215. 215.A. Rahmattalabi, P. Vayanos, K. Dullerud, and E. Rice. Learning resource allocation policies from observational data with an application to homeless services delivery. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 1240–1256, 2022.
  216. 216.I. D. Raji. The anatomy of ai audits: Form, process, and consequences. 2022.
  217. 217.I. D. Raji and L. Liu. Evaluating prediction-based interventions with human decision makers in mind. arXiv preprint arXiv:2503.05704, 2025.
  218. 218.I. D. Raji and L. T. Liu. Designing experimental evaluations of algorithmic interventions with human decision makers in mind. Manuscript, in submission. A preliminary version was presented as a poster at the ICML 2024 workshop on Humans, Algorithmic Decision-Making and Society., 2024.
  219. 219.I. D. Raji, A. Smart, R. N. White, M. Mitchell, T. Gebru, B. Hutchinson, J. Smith-Loud, D. Theron, and P. Barnes. Closing the ai accountability gap: Defining an end-to-end framework for internal algorithmic auditing. In Proceedings of the 2020 conference on fairness, accountability, and transparency, pages 33–44, 2020.
  220. 220.I. D. Raji, I. E. Kumar, A. Horowitz, and A. Selbst. The fallacy of ai functionality. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 959–972, 2022a.
  221. 221.I. D. Raji, P. Xu, C. Honigsberg, and D. Ho. Outsider oversight: Designing a third party audit ecosystem for ai governance. In Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society, pages 557–571, 2022b.
  222. 222.N. E. Reichman, J. O. Teitler, I. Garfinkel, and S. S. McLanahan. Fragile families: Sample and design. Children and Youth Services Review, 23(4-5):303–326, 2001.
  223. 223.E. Rice. The tay triage tool: A tool to identify homeless transition age youth most in need of permanent supportive housing. 2013.
  224. 224.E. Rice, M. Holguin, H.-T. Hsu, M. Morton, P. Vayanos, M. Tambe, and H. Chan. Linking homelessness vulnerability assessments to housing placements and outcomes for youth. Cityscape, 20(3):69–86, 2018.
  225. 225.R. Richardson. Defining and demystifying automated decision systems. Md. L. Rev., 81:785, 2021.
  226. 226.D. G. Robinson and L. Koepke. Civil rights and pretrial risk assessment instruments. Report, page 16, 2019. URL https://safetyandjusticechallenge.org/wp-content/uploads/2021/06/Robinson-Koepke-Civil-Rights-Critical-Issue-Brief.pdf.
  227. 227.A. Rona-Tas. The Off-Label Use of Consumer Credit Ratings. 42:5276, 2017. ISSN 0172-6404. doi: 10.12759/HSR.42.2017.1.52-76. URL http://www.ssoar.info/ssoar/handle/document/51162.
  228. 228.D. Rossman, M. Kurzweil, and B. Lewis. MAAPS Advising Experiment: Evaluation Findings after Six Years. Technical report, Ithaka S+R, May 2023. URL http://sr.ithaka.org/?p=318895.
  229. 229.D. B. Rubin. Causal inference using potential outcomes: Design, modeling, decisions. Journal of the American Statistical Association, 100(469):322–331, 2005.
  230. 230.R. Sahoo and S. Wager. Policy learning with competing agents. arXiv preprint arXiv:2204.01884, 2022.
  231. 231.M. J. Salganik, I. Lundberg, A. T. Kindel, C. E. Ahearn, K. Al-Ghoneim, A. Almaatouq, D. M. Altschul, J. E. Brand, N. B. Carnegie, R. J. Compton, et al. Measuring the predictability of life outcomes with a scientific mass collaboration. Proceedings of the National Academy of Sciences, 117(15):8398–8403, 2020.
  232. 232.A. Sanchez-Becerra. Robust inference for the treatment effect variance in experiments using machine learning. arXiv preprint arXiv:2306.03363, 2023.
  233. 233.D. Saxena and S. Guha. Algorithmic harms in child welfare: Uncertainties in practice, organization, and street-level decision-making. ACM Journal on Responsible Computing, 1(1):1–32, 2024.
  234. 234.D. Saxena, K. Badillo-Urquiola, P. J. Wisniewski, and S. Guha. A framework of high-stakes algorithmic decision-making for the public sector developed through a case study of child-welfare. Proc. ACM Hum.-Comput. Interact., 5(CSCW2), Oct. 2021a. doi: 10.1145/3476089. URL https://doi.org/10.1145/3476089.
  235. 235.D. Saxena, K. Badillo-Urquiola, P. J. Wisniewski, and S. Guha. A framework of high-stakes algorithmic decision-making for the public sector developed through a case study of child-welfare. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW2):1–41, 2021b.
  236. 236.D. Saxena, E. S.-Y. Moon, A. Chaurasia, Y. Guan, and S. Guha. Rethinking ”risk” in algorithmic systems through a computational narrative analysis of casenotes in child-welfare. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY, USA, 2023. Association for Computing Machinery. ISBN 9781450394215. doi: 10.1145/3544548.3581308. URL https://doi.org/10.1145/3544548.3581308.
  237. 237.K. Schechtman, B. Brandon, J. Stafford, H. Li, and L. T. Liu. Discretion in the loop: Human expertise in algorithm-assisted college advising, 2025. URL https://arxiv.org/abs/2505.13325.
  238. 238.J. Schoeffer, M. De-Arteaga, and J. Elmer. Perils of label indeterminacy: A case study on prediction of neurological recovery after cardiac arrest. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’25, page 1080–1094, New York, NY, USA, 2025. Association for Computing Machinery. ISBN 9798400714825. doi: 10.1145/3715275.3732070. URL https://doi.org/10.1145/3715275.3732070.
  239. 239.M. P. Sendak, W. Ratliff, D. Sarro, E. Alderton, J. Futoma, M. Gao, M. Nichols, M. Revoir, F. Yashar, C. Miller, et al. Real-world integration of a sepsis deep learning technology into routine clinical care: implementation study. JMIR medical informatics, 8(7):e15182, 2020.
  240. 240.S. Sirota, N. B. Allen, R. G. Barr, and D. Malinsky. A health equity perspective on data-driven treatment decisions in cardiovascular care: risk assessments versus individualized treatment rules, 2025. Submitted manuscript.
  241. 241.M. Skemer and E. Brennan. Comparing pretrial supervision modes. 2024.
  242. 242.J. A. Smith. Is the threat of reemployment services more effective than the services themselves? experimental evidence from the ui system. 2002.
  243. 243.M. T. Stevenson. Cause, effect, and the structure of the social world. Available at SSRN 4445710, 2023.
  244. 244.G. Stoddard, D. J. Fitzpatrick, and J. Ludwig. Predicting police misconduct. Technical report, National Bureau of Economic Research, 2024.
  245. 245.P. Stone. Why lotteries are just. Journal of political philosophy, 15(3), 2007.
  246. 246.A. Subbaswamy, B. Chen, and S. Saria. A unifying causal framework for analyzing dataset shift-stable learning algorithms. Journal of Causal Inference, 10(1):64–89, 2022.
  247. 247.V. M. Suriyakumar, M. Ghassemi, and B. Ustun. When personalization harms performance: reconsidering the use of group attributes in prediction. In International Conference on Machine Learning, pages 33209–33228. PMLR, 2023.
  248. 248.B. Tang, Ç. Koçyigit, E. Rice, and P. Vayanos. Learning optimal and fair policies for online allocation of scarce societal resources from data collected in deployment. arXiv preprint arXiv:2311.13765, 2023.
  249. 249.L. Temkin. Equality, priority, and the levelling-down objection. 2000.
  250. 250.C. Thompson. Who’s homeless enough for housing? In San Francisco an algorithm decides - Coda Story — codastory.com. https://www.codastory.com/authoritarian-tech/san-francisco-homeless-algorithm/. [Accessed 27-05-2025].
  251. 251.J. Thompson. Mental models and interpretability in ai fairness tools and code environments. In HCI International 2021-Late Breaking Papers: Multimodality, eXtended Reality, and Artificial Intelligence: 23rd HCI International Conference, HCII 2021, Virtual Event, July 24–29, 2021, Proceedings 23, pages 574–585. Springer, 2021.
  252. 252.U.S. Department of Health and Human Services. Vaccine adverse event reporting system. https://vaers.hhs.gov/index.html.
  253. 253.B. Ustun and C. Rudin. Supersparse linear integer models for optimized medical scoring systems. Machine Learning, 102(3):349–391, Nov 2015. ISSN 1573-0565. doi: 10.1007/s10994-015-5528-6. URL http://dx.doi.org/10.1007/s10994-015-5528-6.
  254. 254.B. Ustun and C. Rudin. Learning optimized risk scores. Journal of Machine Learning Research, 2019.
  255. 255.B. Valan, A. Prakash, W. Ratliff, M. Gao, S. Muthya, A. Thomas, J. L. Eaton, M. Gardner, M. Nichols, M. Revoir, et al. Evaluating sepsis watch generalizability through multisite external validation of a sepsis machine learning model. npj Digital Medicine, 8(1):1–10, 2025.
  256. 256.M. Van Smeden, G. Heinze, B. Van Calster, F. W. Asselbergs, P. E. Vardas, N. Bruining, P. De Jaegere, J. H. Moore, S. Denaxas, A. L. Boulesteix, et al. Critical appraisal of artificial intelligence-based prediction models for cardiovascular disease. European Heart Journal, 43(31):2921–2930, 2022.
  257. 257.S. Wachter, B. Mittelstadt, and C. Russell. Bias preservation in machine learning: the legality of fairness metrics under eu non-discrimination law. W. Va. L. Rev., 123:735, 2020.
  258. 258.A. Wang, S. Kapoor, S. Barocas, and A. Narayanan. Against predictive optimization: On the legitimacy of decision-making algorithms that optimize predictive accuracy. ACM Journal on Responsible Computing, 1(1):1–45, 2024a.
  259. 259.S. Wang, S. Bates, P. Aronow, and M. Jordan. On counterfactual metrics for social welfare: Incentives, ranking, and information asymmetry. In International Conference on Artificial Intelligence and Statistics, pages 1522–1530. PMLR, 2024b.
  260. 260.J. R. Wells and B. Weinstock. Pymetrics: International expansion. Harvard Business School Case 720-376, Revised May 2021, 2019. Available at: https://www.hbs.edu/faculty/Pages/item.aspx?num=56675.
  261. 261.M. Wensing. Implementation science in healthcare: Introduction and perspective. Zeitschrift für Evidenz, Fortbildung und Qualität im Gesundheitswesen, 109(2):97–102, 2015.
  262. 262.D. B. White and D. C. Angus. A proposed lottery system to allocate scarce covid-19 medications: promoting fairness and generating knowledge. Jama, 324(4):329–330, 2020.
  263. 263.Wisconsin Department of Public Instruction. Dropout early warning system (dews), 2025. URL https://dpi.wi.gov/ews/dropout. Accessed: 2025-02-25.
  264. 264.A. Wong, E. Otles, J. P. Donnelly, A. Krumm, J. McCullough, O. DeTroyer-Cooley, J. Pestrue, M. Phillips, J. Konye, C. Penoza, et al. External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Internal Medicine, 181(8):1065–1070, 2021.
  265. 265.L. Wynants, B. Van Calster, G. S. Collins, R. D. Riley, G. Heinze, E. Schuit, E. Albu, B. Arshi, V. Bellou, M. M. Bonten, et al. Prediction models for diagnosis and prognosis of covid-19: systematic review and critical appraisal. BMJ, 369, 2020.
  266. 266.A. Xiang. Reconciling legal and technical approaches to algorithmic bias. Tenn. L. Rev., 88:649, 2020.
  267. 267.A. Xiang and I. D. Raji. On the legal compatibility of fairness definitions. arXiv preprint arXiv:1912.00761, 2019.
  268. 268.M. Yurrita, D. Murray-Rust, A. Balayn, and A. Bozzon. Towards a multi-stakeholder value-based assessment framework for algorithmic systems. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’22, pages 535–563, New York, NY, USA, June 2022. Association for Computing Machinery. ISBN 978-1-4503-9352-2. doi: 10.1145/3531146.3533118. URL https://dl.acm.org/doi/10.1145/3531146.3533118.
  269. 269.M. Yurrita, T. Draws, A. Balayn, D. Murray-Rust, N. Tintarev, and A. Bozzon. Disentangling fairness perceptions in algorithmic decision-making: the effects of explanations, human oversight, and contestability. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY, USA, 2023. Association for Computing Machinery. doi: 10.1145/3544548.3581161. URL https://doi.org/10.1145/3544548.3581161.
  270. 270.B. Zacka. When the state meets the street: public service and moral agency. The Belknap Press of Harvard University Press, Cambridge, Massachusetts, 2017. ISBN 978-0-674-54554-0.
  271. 271.Y. Zhang, E. Ben-Michael, and K. Imai. Safe policy learning under regression discontinuity designs. arXiv preprint arXiv:2208.13323, 2022.
  272. 272.Q. Zhao and T. Hastie. Causal interpretations of black-box models. Journal of Business & Economic Statistics, 39(1):272–281, 2021.
  273. 273.S. Zhao, M. Kim, R. Sahoo, T. Ma, and S. Ermon. Calibrating predictions to decisions: A novel approach to multi-class calibration. Advances in Neural Information Processing Systems, 34:22313–22324, 2021.
  274. 274.Y. Zhao, D. Zeng, A. J. Rush, and M. R. Kosorok. Estimating individualized treatment rules using outcome weighted learning. Journal of the American Statistical Association, 107(499):1106–1118, 2012.
  275. 275.A. Zhou. Optimal and fair encouragement policy evaluation and learning, 2023.
  276. 276.A. Zhou, A. Koo, N. Kallus, R. Ropac, R. Peterson, S. Koppel, and T. Bergin. An empirical evaluation of the impact of new york’s bail reform on crime using synthetic controls. arXiv preprint arXiv:2111.08664, 2021.
  277. 277.M. Zilka, R. Fogliato, J. Hron, B. Butcher, C. Ashurst, and A. Weller. The progression of disparities within the criminal justice system: Differential enforcement and risk assessment instruments. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, pages 1553–1569, 2023.

Citation

MLA
Liu, L. T., et al. “Bridging Prediction and Intervention Problems in Social Systems”. arXiv, 2025, http://arxiv.org/abs/2507.05216v3.
APA
Liu, L. T., Raji, I. D., Zhou, A., Guerdan, L., Hullman, J., Malinsky, D., Wilder, B., Zhang, S., Adam, H., Coston, A., Laufer, B., Nwankwo, E., Zanger-Tishler, M., Ben-Michael, E., Barocas, S., Feller, A., Gerchick, M., Gillis, T., Guha, S., … Wilson, A. (2025). Bridging Prediction and Intervention Problems in Social Systems. arXiv. http://arxiv.org/abs/2507.05216v3
Chicago
Liu, L. T., I. D. Raji, A. Zhou, et al. 2025. “Bridging Prediction and Intervention Problems in Social Systems”. arXiv. http://arxiv.org/abs/2507.05216v3.
Harvard
Liu, L.T. et al. (2025) “Bridging Prediction and Intervention Problems in Social Systems”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2507.05216v3.
Vancouver
1. Liu LT, Raji ID, Zhou A, et al (2025) Bridging Prediction and Intervention Problems in Social Systems. arXiv

BibTeX

@article{liu2025bridging,
  title = {Bridging Prediction and Intervention Problems in Social Systems},
  author = {Liu, Lydia T. and Raji, Inioluwa Deborah and Zhou, Angela and Guerdan, Luke and Hullman, Jessica and Malinsky, Daniel and Wilder, Bryan and Zhang, Simone and Adam, Hammaad and Coston, Amanda and Laufer, Ben and Nwankwo, Ezinne and Zanger-Tishler, Michael and Ben-Michael, Eli and Barocas, Solon and Feller, Avi and Gerchick, Marissa and Gillis, Talia and Guha, Shion and Ho, Daniel and Hu, Lily and Imai, Kosuke and Kapoor, Sayash and Loftus, Joshua and Nabi, Razieh and Narayanan, Arvind and Recht, Ben and Perdomo, Juan Carlos and Salganik, Matthew and Sendak, Mark and Tolbert, Alexander and Ustun, Berk and Venkatasubramanian, Suresh and Wang, Angelina and Wilson, Ashia},
  year = {2025},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2507.05216v3},
  eprint = {2507.05216}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/