Semantics derived automatically from language corpora contain human-like biases

Aylin CaliskanJoanna J. BrysonArvind Narayanan

article2016Science3,262 citationsFrontiers of Science Award

Establishes that standard word embeddings automatically acquire human racial, gender, and social prejudices from ordinary text, introducing the Word Embedding Association Test to measure implicit bias in machine learning models.

Listen

This early research demonstrates that widely used machine-learning systems for processing language automatically absorb the same semantic associations and prejudices that humans exhibit. The work addresses growing concerns that artificial intelligence could embed and amplify historic biases in areas such as hiring, criminal justice, and content moderation, even when developers intend no such outcome. Because language models are trained on ordinary web text that reflects centuries of cultural patterns, bias cannot be removed simply by greater transparency, developer diversity, or oversight of individual algorithms.

The authors set out to test whether a standard statistical word-embedding model, trained on a large web corpus, would reproduce well-documented human biases measured by the Implicit Association Test and by real-world employment and naming data. They developed two new evaluation methods, the Word Embedding Association Test and the Word Embedding Factual Association Test, and applied them to the GloVe embedding trained on 840 billion tokens of web text.

The analysis replicated every human bias examined. It recovered the expected pleasantness associations for flowers versus insects and musical instruments versus weapons, with large effect sizes. It reproduced strong racial associations, showing European-American names more closely linked to pleasant terms than African-American names. The same embedding predicted the probability that a résumé with a given name would receive an interview invitation, matching the 50 percent advantage for European-American names found in a large field experiment. Gender associations likewise matched psychological findings: female terms were more strongly linked to family and arts, male terms to career and science. Finally, the model recovered actual 2015 U.S. labor-force participation rates for fifty occupations with a correlation of 0.90 and recovered the gender distribution of common androgynous names with a correlation of 0.84.

These results indicate that prejudice is not an incidental flaw of particular training sets or algorithms but an inherent consequence of learning regularities from human language. Any system that must understand or generate language will therefore carry forward both morally neutral associations and harmful stereotypes unless deliberate countermeasures are introduced after the initial training stage. The findings imply that current calls for algorithmic transparency or more diverse engineering teams, while valuable, are insufficient on their own.

The authors recommend that organizations using language models select training corpora with the least prejudicial content possible, supplement purely statistical representations with explicit symbolic rules or human-curated constraints, and apply the new association tests during development to surface biases before deployment. They note that further interdisciplinary work is required to determine which biases should be mitigated in specific applications and how to do so without destroying useful factual information also encoded in language.

The study relies on a single embedding and corpus; results could differ with other data sources or more recent models. The reported statistical measures apply to word associations rather than to human subjects, so direct numerical comparison with Implicit Association Test effect sizes is not possible. Nonetheless, the consistency of findings across neutral, prejudicial, and veridical associations provides strong evidence that language itself transmits recoverable cultural bias to any system trained on it.

arXiv: 1608.07187
  • Paper: GloVe: Global Vectors for Word Representation, Jeffrey Pennington et al. (2014). Reading this paper first is essential because the source study builds directly upon the GloVe word embedding model introduced here to extract and measure human-like semantic biases.
  • Paper: Efficient Estimation of Word Representations in Vector Space, Tomáš Mikolov et al. (2013). Understanding word2vec is a prerequisite since the source study relies on standard vector representations and distributional semantics to demonstrate how machine learning inherits historical stereotypes.
Cover for Semantics derived automatically from language corpora contain human-like biases

Abstract

Artificial intelligence and machine learning are in a period of astounding growth. However, there are concerns that these technologies may be used, either with or without intention, to perpetuate the prejudice and unfairness that unfortunately characterizes many human institutions. Here we show for the first time that human-like semantic biases result from the application of standard machine learning to ordinary language---the same sort of language humans are exposed to every day. We replicate a spectrum of standard human biases as exposed by the Implicit Association Test and other well-known psychological studies. We replicate these using a widely used, purely statistical machine-learning model---namely, the GloVe word embedding---trained on a corpus of text from the Web. Our results indicate that language itself contains recoverable and accurate imprints of our historic biases, whether these are morally neutral as towards insects or flowers, problematic as towards race or gender, or even simply veridical, reflecting the {\em status quo} for the distribution of gender with respect to careers or first names. These regularities are captured by machine learning along with the rest of semantics. In addition to our empirical findings concerning language, we also contribute new methods for evaluating bias in text, the Word Embedding Association Test (WEAT) and the Word Embedding Factual Association Test (WEFAT). Our results have implications not only for AI and machine learning, but also for the fields of psychology, sociology, and human ethics, since they raise the possibility that mere exposure to everyday language can account for the biases we replicate here.

Table of Contents

  • Introduction
  • Meaning and Bias in Humans and Machines
  • Results
  • Baseline: Replication of Associations That Are Universally Accepted
  • Flowers and Insects
  • **Musical Instruments and Weapons**
  • **Racial Biases**
  • **Replicating Implicit Associations for Valence**
  • *Replicating the Bertrand and Mullainathan (2004) Résumé Study*
  • Gender Biases
  • *Replicating Implicit Associations for Career and Family*
  • **Replicating Implicit Associations for Arts and Mathematics**
  • **Replicating Implicit Associations for Arts and Sciences**
  • **Comparison to Real-World Data: Occupational Statistics**
  • Methods
  • Data and training
  • Word Embedding Association Test (WEAT)
  • Word Embedding Factual Association Test (WEFAT)
  • Discussion
  • Implications for understanding human prejudice
  • Consequences of bias in humans and machines
  • Effects of bias in NLP applications
  • Challenges in addressing bias
  • Awareness is better than blindness
  • Acknowledgements
  • References

Knowls

  1. Knowl 1 — Word Embedding Association Test (WEAT)

    model/method

    The Word Embedding Association Test (WEAT) is a statistical permutation test designed to measure implicit semantic associations and biases between two target concept sets and two attribute sets within vector space word embeddings.

    Let XX and YY be two equal-sized sets of target word vectors (X=Y|X| = |Y|), and let AA and BB be two sets of attribute word vectors. The cosine similarity between two unit-normalized word vectors u\vec{u} and v\vec{v} is denoted by cos(u,v)=uvuv\cos(\vec{u}, \vec{v}) = \frac{\vec{u} \cdot \vec{v}}{\|\vec{u}\| \|\vec{v}\|}.

    The association of a target word vector w\vec{w} with the attribute sets AA and BB is defined as:

    s(w,A,B)=1AaAcos(w,a)1BbBcos(w,b)s(w, A, B) = \frac{1}{|A|} \sum_{a \in A} \cos(\vec{w}, \vec{a}) - \frac{1}{|B|} \sum_{b \in B} \cos(\vec{w}, \vec{b})

    The test statistic s(X,Y,A,B)s(X, Y, A, B) measures the differential association of the target sets XX and YY with the attributes AA and BB:

    s(X,Y,A,B)=xXs(x,A,B)yYs(y,A,B)s(X, Y, A, B) = \sum_{x \in X} s(x, A, B) - \sum_{y \in Y} s(y, A, B)

    The one-sided permutation pp-value calculates the likelihood that the observed difference arises under the null hypothesis of no differential association, computed across all equal-sized partitions {(Xi,Yi)}\{(X_i, Y_i)\} of XYX \cup Y:

    p=Pri[s(Xi,Yi,A,B)>s(X,Y,A,B)]p = \Pr_i\left[s(X_i, Y_i, A, B) > s(X, Y, A, B)\right]

    The effect size dd is the standardized mean difference (analogous to Cohen's dd):

    d=meanxXs(x,A,B)meanyYs(y,A,B)std-devwXYs(w,A,B)d = \frac{\text{mean}_{x \in X} s(x, A, B) - \text{mean}_{y \in Y} s(y, A, B)}{\text{std-dev}_{w \in X \cup Y} s(w, A, B)}

  2. Knowl 2 — Word Embedding Factual Association Test (WEFAT)

    model/method

    The Word Embedding Factual Association Test (WEFAT) evaluates whether word vectors embed continuous, real-world factual properties by examining the correlation between vector-space attribute associations and external empirical ground-truth data.

    Let WW be a set of target word vectors, and let AA and BB be two sets of attribute word vectors. Let pwRp_w \in \mathbb{R} denote an external real-valued property associated with each target word wWw \in W (e.g., the proportion of women employed in occupation ww).

    The normalized association score of a target word vector w\vec{w} with the attribute sets AA and BB is:

    s(w,A,B)=meanaAcos(w,a)meanbBcos(w,b)std-devxABcos(w,x)s(w, A, B) = \frac{\text{mean}_{a \in A} \cos(\vec{w}, \vec{a}) - \text{mean}_{b \in B} \cos(\vec{w}, \vec{b})}{\text{std-dev}_{x \in A \cup B} \cos(\vec{w}, \vec{x})}

    where cos(u,v)\cos(\vec{u}, \vec{v}) is the cosine similarity between vectors u\vec{u} and v\vec{v}.

    The null hypothesis posits that there is no relationship between s(w,A,B)s(w, A, B) and the factual property pwp_w. The hypothesis is evaluated using linear regression or Pearson/Spearman correlation between the embedding association scores {s(w,A,B)}wW\{s(w, A, B)\}_{w \in W} and the empirical values {pw}wW\{p_w\}_{w \in W}.

  3. Knowl 3 — Centroid-Based Filtering of Polysemous Word Vectors in Name Corpora

    algorithm

    When analyzing semantic associations of personal names, static word embeddings conflate proper names that share spelling with common English words (e.g., "Will", "Faith") into a single polysemous vector. To isolate vectors that predominantly reflect name semantics, an algorithmic outlier-rejection procedure filters words whose vectors deviate substantially from the centroid of all candidate name vectors.

    Input: Candidate name set N={w1,w2,,wn}N = \{w_1, w_2, \dots, w_n\}, word embedding function V:wRdV: w \mapsto \mathbb{R}^d, rejection fraction α[0,1)\alpha \in [0, 1) (set to 0.20)
    Output: Filtered name set NfilteredN_{\text{filtered}}
    Compute the embedding vectors for all names: E{V(w):wN}E \leftarrow \{V(w) : w \in N\}
    Compute the name centroid vector: C1EvEvC \leftarrow \frac{1}{|E|} \sum_{\vec{v} \in E} \vec{v}
    For each name wNw \in N:
        Compute distance to centroid: dwV(w)C2d_w \leftarrow \|V(w) - C\|_2
    Rank candidate names in NN in ascending order of dwd_w
    Retain the top (1α)N(1 - \alpha) \cdot |N| names closest to the centroid CC
    return NfilteredN_{\text{filtered}}
  4. Knowl 4 — Correlation Between Word Embedding Gender Associations and Occupational Labor Statistics

    empirical result

    Applying the Word Embedding Factual Association Test (WEFAT) to 300-dimensional GloVe vectors trained on the 840-billion-token Common Crawl corpus demonstrates that embedding associations accurately encode real-world gender distributions across occupations.

    Using the 50 most frequent single-word occupation titles from the 2015 U.S. Bureau of Labor Statistics (such as engineer, mechanic, carpenter, nurse, receptionist, librarian) alongside male attribute terms (male, man, boy, brother, he, him, his, son) and female attribute terms (female, woman, girl, sister, she, her, hers, daughter), the correlation between the embedding-derived female association score and the empirical percentage of women in the occupation is:

    ρ=0.90(p<1018)\rho = 0.90 \quad (p < 10^{-18})

    Furthermore, comparing occupation-gender association scores across different embedding models yields strong agreement: the correlation between GloVe (Common Crawl) and word2vec (Google News corpus) is Pearson ρ=0.88\rho = 0.88 (Spearman ρ=0.86\rho = 0.86).

  5. Knowl 5 — Correlation Between Word Embedding Gender Associations and Census Name Demographics

    empirical result

    Applying the Word Embedding Factual Association Test (WEFAT) to 50 popular androgynous first names (such as Kelly, Tracy, Jamie, Alexis, Dana, Morgan, Robin, Casey) demonstrates that word embeddings mirror demographic name-gender frequencies.

    Using 1990 U.S. Census data for the percentage of individuals with each name who are women, and male vs. female attribute word sets, the correlation between the WEFAT female association scores (after filtering out the 20% least name-like polysemous words via centroid distance) and the actual census gender percentages is:

    ρ=0.84(p<1013)\rho = 0.84 \quad (p < 10^{-13})

  6. Knowl 6 — Replication of Racial Name Valence Biases and Résumé Audit Disparities

    empirical result

    Evaluating pre-trained GloVe embeddings on the Common Crawl corpus using the Word Embedding Association Test (WEAT) replicates human racial biases and employment discrimination studies:

    1. Implicit Association Test (IAT) Racial Bias Replication: Evaluating European American names (Adam, Harry, Josh, Brad, Katie, Meredith, etc.) versus African American names (Alonzo, Jamal, Tyrone, Ebony, Latonya, etc.) against Pleasant words (caress, freedom, peace, love, happy, etc.) and Unpleasant words (abuse, crash, filth, murder, death, etc.) produces a significant preference for European American names with effect size d=1.41d = 1.41 (p<108p < 10^{-8}), replicating the human IAT finding of d=1.17d = 1.17 (p=106p = 10^{-6}).

    2. Résumé Audit Study Replication: Using name sets from the Bertrand and Mullainathan résumé field study on hiring disparities, European American names are significantly closer to Pleasant attributes than African American names. Using original IAT valence attributes yields d=1.50d = 1.50 (p<104p < 10^{-4}); using updated shorter valence attributes (joy, love, peace vs. agony, terrible, horrible) yields d=1.28d = 1.28 (p<103p < 10^{-3}).

  7. Knowl 7 — Replication of Implicit Gender Stereotypes in Career, Mathematics, and Science

    empirical result

    Applying the Word Embedding Association Test (WEAT) to GloVe embeddings reproduces human implicit gender associations documented in psychological literature:

    1. Career vs. Family: Male names (John, Paul, Mike, Kevin, Steve, Greg, Jeff, Bill) versus female names (Amy, Joan, Lisa, Sarah, Diana, Kate, Ann, Donna) evaluated against career words (executive, management, professional, corporation, salary, office, business, career) versus family words (home, parents, children, family, cousins, marriage, wedding, relatives) yields an effect size of d=1.81d = 1.81 (p<103p < 10^{-3}), compared to human online IAT results of d=0.72d = 0.72 (p<102p < 10^{-2}).

    2. Mathematics vs. Arts: Male attributes (male, man, boy, brother, he, him, his, son) versus female attributes (female, woman, girl, sister, she, her, hers, daughter) evaluated against math words (math, algebra, geometry, calculus, etc.) versus arts words (poetry, art, dance, literature, novel, etc.) yields an effect size of d=1.06d = 1.06 (p=102p = 10^{-2}), compared to human IAT results of d=0.82d = 0.82 (p<102p < 10^{-2}).

    3. Science vs. Arts: Male attributes versus female attributes evaluated against science words (science, technology, physics, chemistry, Einstein, NASA, etc.) versus arts words (poetry, art, Shakespeare, dance, literature, etc.) yields an effect size of d=1.24d = 1.24 (p=102p = 10^{-2}), compared to human IAT results of d=1.47d = 1.47 (p=1024p = 10^{-24}).

  8. Knowl 8 — Replication of Baseline Valence Biases for Morally Neutral Categories

    empirical result

    Applying WEAT to GloVe embeddings on morally neutral, universal associations establishes that word co-occurrence statistics alone capture human-like semantic valence without direct physical embodiment:

    • Flowers vs. Insects: 25 flower names (aster, clover, rose, tulip, etc.) versus 25 insect names (ant, flea, spider, wasp, etc.) evaluated against 25 pleasant words (love, peace, happy, etc.) versus 25 unpleasant words (murder, sickness, death, etc.) yields an effect size of d=1.50d = 1.50 (p<107p < 10^{-7}), replicating the original human Implicit Association Test (IAT) finding (d=1.35,p=108d = 1.35, p = 10^{-8}).
    • Musical Instruments vs. Weapons: 25 musical instruments (guitar, violin, piano, flute, etc.) versus 25 weapons (gun, sword, missile, knife, etc.) evaluated against the same pleasant and unpleasant sets yields an effect size of d=1.53d = 1.53 (p<107p < 10^{-7}), replicating the human IAT finding (d=1.66,p=1010d = 1.66, p = 10^{-10}).
  9. Knowl 9 — Gender Stereotype Propagation in Statistical Machine Translation of Gender-Neutral Languages

    empirical result

    Statistical machine translation (SMT) systems translate genderless pronouns from gender-neutral languages (such as Turkish, Finnish, Estonian, Hungarian, and Persian) into gender-stereotyped English pronouns based on word embedding associations.

    In Turkish, which uses the gender-neutral third-person pronoun "O":

    • "O bir doktor." is translated by translation engines to "He is a doctor."
    • "O bir hemşire." is translated to "She is a nurse."

    Across 50 standard occupation words, the translated pronoun is rendered as masculine ("he") in the majority of instances and feminine ("she") in approximately one-quarter of instances. The gender association score derived from word embeddings almost perfectly predicts which gendered pronoun is selected in translation.

  10. Knowl 10 — Conceptual Distinction Between Statistical Semantic Bias and Prejudicial Harm

    theoretical result

    Statistical bias in word representations is indistinguishable from semantic knowledge: embedding algorithms derive meaning by encoding historical and cultural co-occurrence regularities, including veridical empirical facts (e.g., occupational distributions).

    1. Prejudice vs. Informative Bias: Prejudice is defined as a domain-specific subset of statistical bias characterized by harmful social consequences. Because societal definitions of fairness evolve and depend on application context, prejudice cannot be defined or eliminated purely algorithmically.
    2. Limitations of Representation Debiasing: Modifying vector representations via projection-based debiasing acts as "fairness through blindness." It removes veridical information necessary for downstream semantic understanding and remains vulnerable to latent bias recovery via proxy variables.
    3. Decoupled Architecture Recommendation: Rather than blinding the representation layer, AI architectures should maintain accurate perceptual representations of cultural regularities while implementing explicit, deliberative decision-making components to prevent the expression of prejudicial actions.

Coverage note — No substantial contributed material was omitted; the extraction covers the WEAT and WEFAT formulations, data filtering algorithms, all empirical baseline and stereotype replications, demographic correlations, downstream NLP impacts, and the conceptual framing of bias versus prejudice.

References

  1. 1.Angwin, J., Larson, J., Mattu, S., and Kirchner, L. (2016). Machine bias: There’s software used across the country to predict future criminals. and it’s biased against blacks. ProPublica, May, 23.
  2. 2.Barocas, S. and Selbst, A. D. (2014). Big data’s disparate impact. California Law Review, 104.
  3. 3.Barr, A. (2015). Google mistakenly tags black people as ‘gorillas,’ showing limits of algorithms. The New York Times.
  4. 4.Barsalou, L. W. (2009). Simulation, situated conceptualization, and prediction. Philosophical Transactions of the Royal Society B: Biological Sciences, 364(1521):1281–1289.
  5. 5.Bear, A. and Rand, D. G. (2016). Intuition, deliberation, and the evolution of cooperation. Proceedings of the National Academy of Sciences, 113(4):936–941.
  6. 6.Bengio, Y., Ducharme, R., Vincent, P., and Jauvin, C. (2003). A neural probabilistic language model. journal of machine learning research, 3(Feb):1137–1155.
  7. 7.Bertrand, M. and Mullainathan, S. (2004). Are emily and greg more employable than lakisha and jamal? a field experiment on labor market discrimination. The American Economic Review, 94(4):991–1013.
  8. 8.Bishop, C. M. (2006). Pattern Recognition and Machine Learning. Springer, London.
  9. 9.Bolukbasi, T., Chang, K.-W., Zou, J., Saligrama, V., and Kalai, A. (2016). Man is to computer programmer as woman is to homemaker? debiasing word embeddings. arXiv preprint arXiv:1607.06520.
  10. 10.Crawford, K. (2016). Artificial intelligence’s white guy problem. The New York Times.
  11. 11.Dessel, A. (2010). Prejudice in schools: Promotion of an inclusive culture and climate. Education and Urban Society, 42(4):407–429.
  12. 12.Dwork, C., Hardt, M., Pitassi, T., Reingold, O., and Zemel, R. (2012). Fairness through awareness. In Proceedings of the 3rd Innovations in Theoretical Computer Science Conference, pages 214–226. ACM.
  13. 13.Feldman, M., Friedler, S. A., Moeller, J., Scheidegger, C., and Venkatasubramanian, S. (2015). Certifying and removing disparate impact. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 259–268. ACM.
  14. 14.Fitch, W. T. (2004). Kin selection and ‘mother tongues’: A neglected component in language evolution. In Oller, D. K. and Griebel, U., editors, Evolution of Communication Systems: A Comparative Approach, pages 275–296. MIT Press, Cambridge, MA.
  15. 15.Greenwald, A. G., McGhee, D. E., and Schwartz, J. L. (1998). Measuring individual differences in implicit cognition: the implicit association test. Journal of personality and social psychology, 74(6):1464.
  16. 16.Greenwald, A. G. and Pettigrew, T. F. (2014). With malice toward none and charity for some: Ingroup favoritism enables discrimination. American Psychologist, 69(7):669.
  17. 17.Hanheide, M., Göbelbecker, M., Horn, G. S., Pronobis, A., Sjöö, K., Aydemir, A., Jensfelt, P., Gretton, C., Dearden, R., Janicek, M., Zender, H., Kruijff, G.-J., Hawes, N., and Wyatt, J. L. (2015). Robot task planning and explanation in open and uncertain worlds. Artificial Intelligence, pages –. in press.
  18. 18.Kiefer, A. K. and Sekaquaptewa, D. (2007). Implicit stereotypes and women’s math performance: How implicit gender-math stereotypes influence women’s susceptibility to stereotype threat. Journal of Experimental Social Psychology, 43(5):825–832.
  19. 19.Kinzler, K. D., Dupoux, E., and Spelke, E. S. (2007). The native language of social cognition. Proceedings of the National Academy of Sciences, 104(30):12577–12580.
  20. 20.Lee, D. J. (2016). Racial bias and the validity of the implicit association test. Technical Report 35, Helsinki, Finland.
  21. 21.Lowe, W. (1997). Meaning and the mental lexicon. In Proceedings of the 15th International Joint Conference on Artificial Intelligence, pages 1092–1097, Nagoya. Morgan Kaufmann.
  22. 22.Macfarlane, T. (2013). Extracting semantics from the enron corpus.
  23. 23.McDonald, S. and Lowe, W. (1998). Modelling functional priming and the associative boost. In Proceedings of the Twentieth Annual Conference of the Cognitive Science Society, pages 667–680. LEA.
  24. 24.Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.
  25. 25.Mikolov, T. and Dean, J. (2013). Distributed representations of words and phrases and their compositionality. Advances in neural information processing systems.
  26. 26.Monteith, L. L. and Pettit, J. W. (2011). Implicit and explicit stigmatizing attitudes and stereotypes about depression. Journal of Social and Clinical Psychology, 30(5):484.
  27. 27.Moss, H. E., Ostrin, R. K., Tyler, L. K., and Marslen-Wilson, W. D. (1995). Accessing different types of lexical semantic information: Evidence from priming. Journal of Experimental Psychology: Learning, Memory and Cognition, 21:863–883.
  28. 28.Noble, S. U. (2013). Google search: Hyper-visibility as a means of rendering black women and girls invisible. InVisible Culture, 19.
  29. 29.Nosek, B. A., Banaji, M., and Greenwald, A. G. (2002a). Harvesting implicit group attitudes and beliefs from a demonstration web site. Group Dynamics: Theory, Research, and Practice, 6(1):101.
  30. 30.Nosek, B. A., Banaji, M. R., and Greenwald, A. G. (2002b). Math= male, me= female, therefore math6= me. Journal of personality and social psychology, 83(1):44.
  31. 31.Nosek, B. A., Smyth, F. L., Sriram, N., Lindner, N. M., Devos, T., Ayala, A., Bar-Anan, Y., Bergh, R., Cai, H., Gonsalkorale, K., et al. (2009). National differences in gender–science stereotypes predict national sex differences in science and math achievement. Proceedings of the National Academy of Sciences, 106(26):10593–10597.
  32. 32.Oswald, M. and Grace, J. (2016). Norman stanley fletcher and the case of the proprietary algorithmic risk assessment. Policing Insight.
  33. 33.Pennington, J., Socher, R., and Manning, C. D. (2014). Glove: Global vectors for word representation. In EMNLP, volume 14, pages 1532–43.
  34. 34.Purcell, B. A. and Kiani, R. (2016). Hierarchical decision processes that operate over distinct timescales underlie choice and changes in strategy. Proceedings of the National Academy of Sciences, 113(31):E4531–E4540.
  35. 35.Quine, W. V. O. (1960). Word and Object. MIT Press, Cambridge, MA.
  36. 36.Silva, A. S. and Mace, R. (2015). Inter-group conflict and cooperation: field experiments before, during and after sectarian riots in northern ireland. Frontiers in Psychology, 6:1790.
  37. 37.Stanley, D. A., Sokol-Hessner, P., Banaji, M. R., and Phelps, E. A. (2011). Implicit race attitudes predict trustworthiness judgments and economic trust decisions. Proceedings of the National Academy of Sciences, 108(19):7710–7715.
  38. 38.Sweeney, L. (2013). Discrimination in online ad delivery. Queue, 11(3):10:10–10:29.
  39. 39.Thórisson, K. R. (2007). Integrated A.I. systems. Minds and Machines, 17(1):11–25.
  40. 40.Turing, A. M. (1950). Computing machinery and intelligence. Mind, 59(236):433–460.
  41. 41.Winner, L. (1980). Do artifacts have politics? Daedalus, pages 121–136.
  42. 42.Zaromb, F. M., Howard, M. W., Dolan, E. D., Sirotin, Y. B., Tully, M., Wingfield, A., and Kahana, M. J. (2006). Temporal associations and prior-list intrusions in free recall. Journal of Experimental Psychology: Learning, Memory, and Cognition, 32(4):792.
  43. 43.Zemel, R. S., Wu, Y., Swersky, K., Pitassi, T., and Dwork, C. (2013). Learning fair representations. ICML (3), 28:325–333.
  44. 44.Zue, V. W. (1985). The use of speech knowledge in automatic speech recognition. Proceedings of the IEEE, 73(11):1602–1615.

Citation

MLA
Caliskan, A., et al. “Semantics Derived Automatically from Language Corpora Contain Human-like Biases”. Science, vol. 356, no. 6334, 2017, pp. 183–86, https://doi.org/10.1126/science.aal4230.
APA
Caliskan, A., Bryson, J. J., & Narayanan, A. (2017). Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334), 183–186. https://doi.org/10.1126/science.aal4230
Chicago
Caliskan, A., J. J. Bryson, and A. Narayanan. 2017. “Semantics Derived Automatically from Language Corpora Contain Human-like Biases”. Science 356 (6334): 183–86. https://doi.org/10.1126/science.aal4230.
Harvard
Caliskan, A., Bryson, J.J. and Narayanan, A. (2017) “Semantics derived automatically from language corpora contain human-like biases”, Science, 356(6334), pp. 183–186. Available at: https://doi.org/10.1126/science.aal4230.
Vancouver
1. Caliskan A, Bryson JJ, Narayanan A (2017) Semantics derived automatically from language corpora contain human-like biases. Science 356:183–186

BibTeX

@article{Caliskan_2017, title={Semantics derived automatically from language corpora contain human-like biases}, volume={356}, ISSN={1095-9203}, url={http://dx.doi.org/10.1126/science.aal4230}, DOI={10.1126/science.aal4230}, number={6334}, journal={Science}, publisher={American Association for the Advancement of Science (AAAS)}, author={Caliskan, Aylin and Bryson, Joanna J. and Narayanan, Arvind}, year={2017}, month=Apr, pages={183–186} }
Metadata:Crossref

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF