Hybrid Uncertainty Quantification for Selective Text Classification in Ambiguous Tasks
Artem Vazhentsev1,2^{1,2}1,2, Gleb Kuzmin1,4^{1,4}1,4, Akim Tsvigun5,8^{5,8}5,8,
Alexander Panchenko
2,1^{2,1}2,1, Maxim Panov6^{6}6, Mikhail Burtsev7^{7}7, and Artem Shelmanov3^{3}3
1^{1}1AIRI, 2^{2}2Skoltech, 3^{3}3MBZUAI, 4^{4}4FRC CSC RAS, 5^{5}5AI Center NUST MISiS, 6^{6}6TII,
7^{7}7London Institute for Mathematical Sciences, 8^{8}8Semrush
{vazhentsev, kuzmin, panchenko}@airi.net
[email protected] [email protected]
Abstract
Many text classification tasks are inherently ambiguous, which results in automatic systems having a high risk of making mistakes, in spite of using advanced machine learning models. For example, toxicity detection in user-generated content is a subjective task, and notions of toxicity can be annotated according to a variety of definitions that can be in conflict with one another. Instead of relying solely on automatic solutions, moderation of the most difficult and ambiguous cases can be delegated to human workers. Potential mistakes in automated classification can be identified by using uncertainty estimation (UE) techniques. Although UE is a rapidly growing field within natural language processing, we find that state-of-the-art UE methods estimate only epistemic uncertainty and show poor performance, or under-perform trivial methods for ambiguous tasks such as toxicity detection. We argue that in order to create robust uncertainty estimation methods for ambiguous tasks it is necessary to account also for aleatoric uncertainty. In this paper, we propose a new uncertainty estimation method that combines epistemic and aleatoric UE methods. We show that by using our hybrid method, we can outperform state-of-the-art UE methods for toxicity detection and other ambiguous text classification tasks1.
1 Introduction
Many natural language processing (NLP) tasks are subjective and contain inherent ambiguity. For example, the notion of toxicity is inherently subjective (Waseem, 2016) and can be defined in a number of ways that may conflict with one another and differ according to the demographic that the methods are applied to (Thylstrup and Waseem, 2020). For many datasets, implicit or ambiguous toxicity can comprise more than 90% of the labeled toxic content (Hartvigsen et al., 2022). Such ambiguity introduces a high risk of classification mistakes for machine learning (ML) models. Classification mistakes for toxicity detection can result in the removal of legitimate non-toxic content on one hand, and the lack of sanction for toxic content, on the other. A common method for addressing this concern for content moderation is to abstain from predictions on ambiguous instances and process them with the help of human workers (Roberts, 2019).
A classification task where some model predictions can be “rejected” is called selective classification (Geifman and El-Yaniv, 2017). The common approach to solving it is applying uncertainty estimation (UE) techniques. UE is a field of ML that seeks to model the degree to which model predictions can be trusted by correlating model mistakes and performance. Better UE methods improve the performance of selective classification and the trade-off between the amount of labor and the reliability of downstream applications. In toxicity detection, better UE methods minimize the amount of content that is reviewed by human moderators to predominately be classification errors.
Recent works have suggested deterministic approaches to UE of neural network predictions based on fitting the density of latent instance representations (Lee et al., 2018; van Amersfoort et al., 2020; Mukhoti et al., 2023; Yoo et al., 2022; Kotelevskii et al., 2022). They have shown good performance in NLP for the detection of out-of-distribution (OOD) instances, adversarial attacks, and misclassified objects in non-ambiguous tasks. However, they primarily capture epistemic uncertainty, i.e. uncertainty related to the lack of knowledge about model parameters and training data, overlooking aleatoric uncertainty, i.e. uncertainty that arises from ambiguity and noise in data.
This work aims to create a UE method for more reliable selective classification in ambiguous tasks such as toxicity detection by combining different types of uncertainty. Instances that carry a high risk of classification mistakes come from two sources: a) OOD areas, which can be detected with epistemic UE methods; and b) in-distribution ambiguous areas, for detection of which, aleatoric UE methods are appropriate (for illustration see Figure 1). Therefore, we propose a Hybrid Uncertainty Quantification (HUQ) method that switches between epistemic and aleatoric uncertainties or linearly combines them. It produces better scores of total uncertainty, which subsequently leads to better selective classification. The experiments on various ambiguous tasks show that HUQ in a majority of cases significantly outperforms other state-of-the-art UE techniques. To summarize, the contributions of this work are the following.
-
In Section 4, we propose a new uncertainty estimation method HUQ that combines epistemic and aleatoric UE techniques in a special way that allows to improve the quality of selective classification in ambiguous tasks.
-
To the best of our knowledge, this work is the first to conduct an empirical investigation of state-of-the-art UE methods for ambiguous text classification tasks such as toxicity detection. Our analysis shows that the proposed HUQ approach outperforms state-of-the-art methods in selective text classification on ambiguous tasks; see Sections 5 and 6.
-
We analyze the limitations of the proposed method and suggest conditions to be met for achieving the improvements; see Section 7.
2 Related Work
Quantifying uncertainty of deep neural network predictions can be successfully accomplished using deep ensembles (DE; Lakshminarayanan et al., 2017), Bayesian models (Blundell et al., 2015), or their approximations. However, most of these methods have various drawbacks, including large computational overhead. For example, for DE, we need to multiply training time, the occupied memory, and inference time, since this network requires training, storing, and running inference for multiple versions of the same model. This makes DE hardly applicable in real-world scenarios.
Recent work has investigated computationally efficient deterministic approaches (e.g., Lee et al., 2018; van Amersfoort et al., 2020; Liu et al., 2020). However, most work is based on feature space density and focuses only on the OOD detection task and epistemic uncertainty estimation. Another computationally efficient approach is SelectiveNet (Geifman and El-Yaniv, 2019), which was designed for computer vision tasks. It introduces two separate heads for prediction and selection within the model architecture and adds a special loss component to minimize selective risk with a specified coverage.
Most similar to our work is Mukhoti et al. (2023), which also considers both aleatoric and epistemic uncertainty. DDU uses a combination of feature-space density for epistemic uncertainty and the softmax predictive distribution for aleatoric uncertainty. They advocate for the usage of different methods for quantifying uncertainty, depending on whether a considered instance is ID or OOD. However, they overlook using a linear combination of uncertainty scores, relying solely on feature-space density for instances considered OOD. We note that these instances can also be borderline (instances from middle to low-density areas), for which using aleatoric uncertainty measures may also be appropriate. Besides, Mukhoti et al. (2023) do not provide results for selective classification and mostly experiment with image classification tasks.
Recently, selective classification (or misclassification detection) has been studied for NLP tasks. One line of such work has proposed adding a regularization term to the training loss. Xin et al. (2021) introduces a penalty term for confident instances with a high loss value. Another approach proposed by Zhang et al. (2019) uses a metric regularization that minimizes the inter-class distance in the latent feature space while maximizing the margin between classes. He et al. (2020) propose a regularization technique based on self-ensembling that aims to minimize the difference between predictions of the two versions of the model. They also combine this approach with mix-up (Thulasidasan et al., 2019) and a distinctiveness score based on the MD. Some work has also considered approximations of deep ensembles based on Monte-Carlo dropout (e.g., Shelmanov et al., 2021; Vazhentsev et al., 2022). Vazhentsev et al. (2022) conduct a vast empirical investigation and suggest several promising combinations of regularizers and feature-density-based methods. They also highlight the importance of spectral normalization for obtaining good results. Kotelevskii et al. (2022) propose a new UE method NUQ and test it for text classification models trained in the low-resource regime.
Despite the aforementioned efforts, highly ambiguous text classification tasks such as toxicity detection have been overlooked in the previous work. Moreover, to the best of our knowledge, no prior work in NLP takes into account aleatoric uncertainty and combines multiple types of uncertainty for a holistic view of uncertainty.
3 Background
Two types of uncertainty have been documented in the literature: aleatoric and epistemic (Der Kiureghian and Ditlevsen, 2009). Aleatoric, or data uncertainty, arises from ambiguity and noise in data. It should be high, for example, for groups of instances prone to annotation discrepancy. Epistemic, or model uncertainty, pertains to a lack of knowledge about model parameters and can often be mitigated through additional training data collection. Epistemic uncertainty is particularly important for OOD detection (Hendrycks and Gimpel, 2017) and active learning (Settles, 2009).
According to the Bayesian approach to measuring uncertainty in deep learning networks (Blundell et al., 2015; Gal, 2016; Depeweg et al., 2018), the total uncertainty of a model prediction x\mathbf{x}x is a sum of aleatoric UA(x)U_{\mathrm{A}}(\mathbf{x})UA(x) and epistemic uncertainty UE(x)U_{\mathrm{E}}(\mathbf{x})UE(x):
High total uncertainty should correlate with classification mistakes and can be used to flag model predictions for human review.
3.1 Out-of-Distribution and (Ambiguous) In-Distribution Instances
We define out-of-distribution (OOD) instances XOOD\mathcal{X}_{\mathrm{OOD}}XOOD as those located either outside a training data distribution or in its low-density regions. They can be identified by high epistemic uncertainty.
In-distribution (ID) instances we define to belong to the domain of the dataset D\mathcal{D}D located “inside” the training data distribution. ID instances are those, for which model predictions have very small epistemic uncertainty, i.e. below some threshold δmin\delta_{\min}δmin:
Note that for the in-distribution data, on the basis of (1) and taking into account (2), we can empirically approximate UT(x)≃UA(x)U_{\mathrm{T}}(\mathbf{x}) \simeq U_{\mathrm{A}}(\mathbf{x})UT(x)≃UA(x).
We also define ambiguous in-distribution (AID) instances as those, predictions on which having the highest values of aleatoric uncertainty with a lower bound δmax\delta_{\max}δmax. AID instances lie around the class-decision boundaries virtually established by the discriminative model:
3.2 Quantifying Epistemic Uncertainty
Recent works have proposed a variety of computationally efficient methods for quantifying epistemic uncertainty on the basis of fitting the probability density of latent instance representations. In this work, we experiment with Mahalanobis Distance (MD, Lee et al., 2018), Robust Density Estimation (RDE, Yoo et al., 2022), and Deep Deterministic Uncertainty (DDU, Mukhoti et al., 2023).
Let D\mathcal{D}D be a training dataset, h(x)h(\mathbf{x})h(x) be a latent representation of an instance x\mathbf{x}x (it is usually taken from the penultimate layer of the network), and c∈Cc \in Cc∈C be a class. The UE method based on MD (Lee et al., 2018), for each class, fits a Gaussian centered in a class centroid {μc}c∈C\{\mu_c\}_{c \in C}{μc}c∈C with a covariance matrix Σ\SigmaΣ shared across classes. The highest class-conditional probability density p(h(x)∣y=c)p(h(\mathbf{x}) \mid y = c)p(h(x)∣y=c) determines the confidence of the prediction, and the uncertainty score is computed as the Mahalanobis distance between h(x)h(\mathbf{x})h(x) and the closest centroid:
RDE (Yoo et al., 2022) improves on MD by computing the covariance matrix Σc\Sigma_cΣc for each individual class using the Minimum Covariance Determinant estimation (Rousseeuw, 1984) and by reducing the dimensionality of the hidden representations via PCA decomposition with an RBF kernel. These modifications aim to minimize the determinant of the covariance matrix and reduce the influence of outliers in the training data.
DDU (Mukhoti et al., 2023) fits a Gaussian Mixture Model (GMM) p(h(x),y)p(h(\mathbf{x}), y)p(h(x),y) with a single mixture component per class. The uncertainty score is the probability density of h(x)h(\mathbf{x})h(x) under the GMM:
where p(h(x)∣y=c)∼N(h(x)∣μc,Σc)p(h(\mathbf{x}) \mid y = c) \sim \mathcal{N}(h(\mathbf{x}) \mid \mu_c, \Sigma_c)p(h(x)∣y=c)∼N(h(x)∣μc,Σc) and p(y=c)=1∣D∣∑(xi,yi)∈D1[yi=c]p(y = c) = \frac{1}{|\mathcal{D}|} \sum_{(\mathbf{x}_i, y_i) \in \mathcal{D}} \mathbf{1}[y_i = c]p(y=c)=∣D∣1∑(xi,yi)∈D1[yi=c].
Methods based on the fitting density of latent representations are suitable for finding OOD instances but are not capable of identifying AID instances. More generally, they are not good estimators of uncertainty in XID\mathcal{X}_{\mathrm{ID}}XID. Therefore, for ambiguous tasks where AID instances comprise a large portion of the data, these epistemic UE methods cannot fully cover all potential misclassifications.
3.3 Quantifying Aleatoric Uncertainty
As measures of aleatoric uncertainty, we use two well-known methods based on probabilities from the output softmax layer of a neural network: entropy (Gal, 2016) and Softmax Response (SR, Geifman and El-Yaniv, 2017):
Entropy and SR have been proposed also as measures of total uncertainty (Malinin and Gales, 2018). However, this assumption holds only when one has access to the full posterior distribution under the Bayesian paradigm, i.e. all possible uncertainties are quantified within the model. In practice, training datasets are limited, and we can only approximate considered probability distributions. Thus, these methods could not capture all the epistemic uncertainty and mostly reflect the aleatoric one (van Amersfoort et al., 2020; Mukhoti et al., 2023).
4 Hybrid Uncertainty Quantification
There are two major sources of mistakes in model predictions: OOD instances and instances that lie in proximity to the decision boundary (AID instances). Aleatoric uncertainty can help to detect AID instances, while epistemic uncertainty can help to detect OOD instances. In many tasks, we have to deal with both types of mistakes arising from task ambiguity or from a marked covariate shift between training and test data. To address this issue, we propose a hybrid method that combines the strengths of aleatoric and epistemic uncertainty.
Our hybrid uncertainty quantification (HUQ) method first uses Eq. (2) to determine whether an instance x\mathbf{x}x is ID or OOD. If x∈XID\mathbf{x} \in \mathcal{X}_{\mathrm{ID}}x∈XID, HUQ applies Eq. (3) to determine if x\mathbf{x}x is near a class-decision boundary, i.e., x∈XAID\mathbf{x} \in \mathcal{X}_{\mathrm{AID}}x∈XAID. Once the type of instance has been identified, we can apply an appropriate uncertainty estimation method for it or combine multiple uncertainty scores into a single estimate.
Uncertainty scores from different methods may however not be comparable with one another due to different magnitudes. Therefore, instead of using absolute values, we propose to rank instances in some dataset D\mathfrak{D}D by their uncertainty scores and as a final score use these ranks or their combinations. Ranking can be considered as a form of normalization. Moreover, such an approach is desirable for the selective classification task, as we are only interested in the ability to rank predictions by their uncertainty. We define a ranking function R(u,D)R(\mathbf{u}, \mathfrak{D})R(u,D) as the rank of u\mathbf{u}u over a sorted dataset D\mathfrak{D}D, so u1>u2\mathbf{u}_1 > \mathbf{u}_2u1>u2 implies R(u1,D)>R(u2,D)R(\mathbf{u}_1, \mathfrak{D}) > R(\mathbf{u}_2, \mathfrak{D})R(u1,D)>R(u2,D).
Having the ranks according to epistemic and aleatoric scores and the type of x\mathbf{x}x, we can define the final total uncertainty score. We consider predictions for ID instances as the most trustworthy, therefore, we define their total uncertainty score as the rank of their aleatoric score R(UA(x),DID)R(U_{\mathrm{A}}(\mathbf{x}), \mathcal{D}_{\mathrm{ID}})R(UA(x),DID) only among known ID instances DID={xi:xi∈D∩xi∈XID}\mathcal{D}_{\mathrm{ID}} = \{\mathbf{x}_i : \mathbf{x}_i \in \mathcal{D} \cap \mathbf{x}_i \in \mathcal{X}_{\mathrm{ID}}\}DID={xi:xi∈D∩xi∈XID}. Predictions on AID instances are considered the most error-prone. Their total score is the rank of the aleatoric score among all known instances R(UA(x),D)R(U_{\mathrm{A}}(\mathbf{x}), \mathcal{D})R(UA(x),D). Lastly, for x∉XID\mathbf{x} \notin \mathcal{X}_{\mathrm{ID}}x∈/XID, we calculate a linear combination of ranks of aleatoric and epistemic scores among all known instances: (1−α)R(UE(x),D)+αR(UA(x),D)(1 - \alpha) R(U_{\mathrm{E}}(\mathbf{x}), \mathcal{D}) + \alpha R(U_{\mathrm{A}}(\mathbf{x}), \mathcal{D})(1−α)R(UE(x),D)+αR(UA(x),D), where α∈[0,1]\alpha \in [0, 1]α∈[0,1] is a task-specific hyperparameter that depends on the quality of the softmax classifier, and a number of training instances. The usage of a mixture rather than only the epistemic score is justified by the fact that the generalization capabilities of models allow them to make meaningful predictions also in OOD regions, so aleatoric scores to some extent remain meaningful in these areas.
Thus, the total uncertainty score for x\mathbf{x}x according to HUQ is
Note that HUQ can plug-in various “base” methods for the estimation of epistemic and aleatoric uncertainty. Algorithm 1 summarizes the uncertainty score calculation procedure according to HUQ.
The threshold hyperparameters (δmin\delta_{\min}δmin, δmax\delta_{\max}δmax) that determine x∈{XID∣XAID∣XOOD}\mathbf{x} \in \{\mathcal{X}_{\mathrm{ID}} \mid \mathcal{X}_{\mathrm{AID}} \mid \mathcal{X}_{\mathrm{OOD}}\}x∈{XID∣XAID∣XOOD} can be set using the validation dataset. We set δmin\delta_{\min}δmin to be the epistemic uncertainty score of the instances x\mathbf{x}x with the lowest β%\beta\%β% epistemic uncertainty on the training set. Similarly, the hyperparameter δmax\delta_{\max}δmax is selected as the uncertainty score of the most confident instances x\mathbf{x}x from top γ%\gamma\%γ% of instances in the training set with the highest aleatoric uncertainty:
5 Experimental Setup
5.1 Models
We experiment with two pre-trained Transformers: ELECTRA (“electra-base-discriminator”) (Clark et al., 2020) and BERT (“bert-base-uncased”) (Devlin et al., 2019) with 110 million parameters. We use a spectral normalization of the weight matrix in the penultimate linear layer of the classification heads of the models (Liu et al., 2020) as it can be helpful for density-based methods (Vazhentsev et al., 2022). The details on the model hyperparameter optimization procedure and optimal values are presented in Appendix A. To report the deviation of results, for each experiment, we train 5 models with optimal hyperparameters, but different random seeds.
5.2 Datasets
There are several tasks that contain highly subjective data, e.g., toxicity detection, particularly detecting implicit hate and sentiment analysis. We conduct experiments on five datasets for toxicity detection: PARADETOX (Logacheva et al., 2022), JIGSAW with binary labels,2 a collection of tweets with annotation of hate and offensive language (TWITTER; Davidson et al., 2017), TOXIGEN (Hartvigsen et al., 2022), and IMPLICITHATE (ElSherief et al., 2021); and three multi-class classification tasks with high ambiguity: 20 NEWS GROUPS (Lang, 1995), Stanford Sentiment Treebank with 5 classes (SST-5; Socher et al., 2013), and AMAZON REVIEWS (McAuley and Leskovec, 2013) (sports and outdoors categories). Note that for TOXIGEN and IMPLICITHATE, implicit hate speech accounts for more than 95% of the positive class. The TWITTER dataset does not contain a predefined test set, so we create it by ourselves. It is constructed from the documents with high annotator disagreement. In all other cases, we use original test sets. See Appendix B for dataset statistics and the analysis of their ambiguity.
To reduce the computational burden of the experiments, the datasets are randomly subsampled. For training, we sample 10% from AMAZON, IMPLICITHATE, and JIGSAW; and 20% from PARADETOX. For evaluation, we sample 10% from PARADETOX, IMPLICITHATE, and JIGSAW.
5.3 Metrics
Selective classification differs from the standard classification task as low certainty predictions are rejected and deferred to alternate procedures, e.g., human review. Therefore, for performance evaluation in this task, a special metric is used: area under the risk coverage curve (AUC-RC; El-Yaniv and Wiener, 2010). Consider all predictions in a dataset are sorted in ascending order by uncertainty, so we can discard some % of the most uncertain predictions. The % of predictions remaining after that is called a coverage rate, and the total loss of the remaining predictions is called the selective risk. The RC curve plots a dependence of the selective risk from the coverage rate. Finally, the AUC-RC is a cumulative sum of the selective losses for each coverage rate. Lower values of AUC-RC indicate better performance.
5.4 Hyperparameter Selection for HUQ
To find optimal hyperparameters for HUQ, we select 20% of the training set as a validation set and optimize AUC-RC on it, using a grid search. For each variant of models trained with different random seeds, we select its specific set of hyperparameters. The hyperparameter grid is the following: α∈[0;1]\alpha \in [0; 1]α∈[0;1] with a step size 0.1; δmin∈{0%,0.05%,0.1%,0.15%,0.2%}\delta_{\min} \in \{0\%, 0.05\%, 0.1\%, 0.15\%, 0.2\%\}δmin∈{0%,0.05%,0.1%,0.15%,0.2%}; δmax∈{0.9%,0.95%,1.0%}\delta_{\max} \in \{0.9\%, 0.95\%, 1.0\%\}δmax∈{0.9%,0.95%,1.0%}. The values of δmax\delta_{\max}δmax and δmin\delta_{\min}δmin in % are converted into absolute values, when we apply them to the test data.
6 Results
In our illustrative example of the two moons dataset in Figure 1, the state-of-the-art epistemic UE methods, MD and RDE, separate the ID area from the remaining feature space well. However, the middle area between the two classes is marked with high confidence, yet for SR and Entropy, this area is marked as highly uncertain due to the presence of instances with high aleatoric uncertainty. HUQ, which combines aleatoric and epistemic uncertainty, accurately detects both areas of uncertainty, thereby overcoming the weaknesses of aleatoric and epistemic uncertainty individually applied.
lt;$UE method
gt;$). Note that in the main part of the paper, we present the results only with SR as a base aleatoric UE method. The results for entropy are very similar to SR and are presented in Appendix E.
HUQ-MD yields significant improvements over its base methods (MD and SR) on 6/8 datasets for both ELECTRA and BERT (see Table 1). The largest improvements are achieved on 20 NEWS GROUPS and PARADETOX, where HUQ reduces AUC-RC by 13.0% and 4.9% (ELECTRA), and 6.6% and 7.0% (BERT).
HUQ-DDU produces improvements over DDU and SR on all 8 datasets for ELECTRA and on 5 datasets for BERT (see Table 2). For ELECTRA, HUQ produces large effects on PARADETOX and TOXIGEN with 4.6% and 11% AUC-RC reduction, and with BERT on PARADETOX with a 10.6% reduction. Interestingly, vanilla DDU is significantly outperformed by SR for ELECTRA on JIGSAW. Applying HUQ addresses this issue, and improves on the results using SR by 1.8%.
The results for HUQ-RDE are more ambiguous than for DDU and SR (see Table 3). RDE is a good method for selective classification and is a hard-to-beat baseline for HUQ. This is because, in addition to OOD detection RDE computes a covariance matrix for each class, thereby making it suitable for identifying decision boundaries. For RDE as the base epistemic UE method, HUQ improves results on 4 datasets for ELECTRA and on 6 datasets for BERT. On some datasets, HUQ does not improve on SR and RDE, e.g., for TWITTER (ELECTRA) and PARADETOX (BERT). However, on others, HUQ shows big improvements in RC-AUC, e.g., 18.0% for 20 NEWS GROUPS (ELECTRA) and 5.0% for TOXIGEN (BERT).
Overall, we see that HUQ usually improves upon its base methods, but in some cases, retains the same performance. We suspect that the configurations where HUQ does not outperform the baselines are due to the presence of large covariate shifts between the training and test data. We discuss this in detail in Section 7.
" data-original-markdown="
### 6.1 HUQ Against its Base Methods
When presenting results, we denote HUQ with a specific base epistemic UE method as HUQ (
lt;$UE method
gt;$). Note that in the main part of the paper, we present the results only with SR as a base aleatoric UE method. The results for entropy are very similar to SR and are presented in Appendix E.
HUQ-MD yields significant improvements over its base methods (MD and SR) on 6/8 datasets for both ELECTRA and BERT (see Table 1). The largest improvements are achieved on 20 NEWS GROUPS and PARADETOX, where HUQ reduces AUC-RC by 13.0% and 4.9% (ELECTRA), and 6.6% and 7.0% (BERT).
HUQ-DDU produces improvements over DDU and SR on all 8 datasets for ELECTRA and on 5 datasets for BERT (see Table 2). For ELECTRA, HUQ produces large effects on PARADETOX and TOXIGEN with 4.6% and 11% AUC-RC reduction, and with BERT on PARADETOX with a 10.6% reduction. Interestingly, vanilla DDU is significantly outperformed by SR for ELECTRA on JIGSAW. Applying HUQ addresses this issue, and improves on the results using SR by 1.8%.
The results for HUQ-RDE are more ambiguous than for DDU and SR (see Table 3). RDE is a good method for selective classification and is a hard-to-beat baseline for HUQ. This is because, in addition to OOD detection RDE computes a covariance matrix for each class, thereby making it suitable for identifying decision boundaries. For RDE as the base epistemic UE method, HUQ improves results on 4 datasets for ELECTRA and on 6 datasets for BERT. On some datasets, HUQ does not improve on SR and RDE, e.g., for TWITTER (ELECTRA) and PARADETOX (BERT). However, on others, HUQ shows big improvements in RC-AUC, e.g., 18.0% for 20 NEWS GROUPS (ELECTRA) and 5.0% for TOXIGEN (BERT).
Overall, we see that HUQ usually improves upon its base methods, but in some cases, retains the same performance. We suspect that the configurations where HUQ does not outperform the baselines are due to the presence of large covariate shifts between the training and test data. We discuss this in detail in Section 7.
" data-source-offset="29531" class="markdown-segment">
6.1 HUQ Against its Base Methods
When presenting results, we denote HUQ with a specific base epistemic UE method as HUQ (<<<UE method>>>). Note that in the main part of the paper, we present the results only with SR as a base aleatoric UE method. The results for entropy are very similar to SR and are presented in Appendix E.
HUQ-MD yields significant improvements over its base methods (MD and SR) on 6/8 datasets for both ELECTRA and BERT (see Table 1). The largest improvements are achieved on 20 NEWS GROUPS and PARADETOX, where HUQ reduces AUC-RC by 13.0% and 4.9% (ELECTRA), and 6.6% and 7.0% (BERT).
HUQ-DDU produces improvements over DDU and SR on all 8 datasets for ELECTRA and on 5 datasets for BERT (see Table 2). For ELECTRA, HUQ produces large effects on PARADETOX and TOXIGEN with 4.6% and 11% AUC-RC reduction, and with BERT on PARADETOX with a 10.6% reduction. Interestingly, vanilla DDU is significantly outperformed by SR for ELECTRA on JIGSAW. Applying HUQ addresses this issue, and improves on the results using SR by 1.8%.
The results for HUQ-RDE are more ambiguous than for DDU and SR (see Table 3). RDE is a good method for selective classification and is a hard-to-beat baseline for HUQ. This is because, in addition to OOD detection RDE computes a covariance matrix for each class, thereby making it suitable for identifying decision boundaries. For RDE as the base epistemic UE method, HUQ improves results on 4 datasets for ELECTRA and on 6 datasets for BERT. On some datasets, HUQ does not improve on SR and RDE, e.g., for TWITTER (ELECTRA) and PARADETOX (BERT). However, on others, HUQ shows big improvements in RC-AUC, e.g., 18.0% for 20 NEWS GROUPS (ELECTRA) and 5.0% for TOXIGEN (BERT).
Overall, we see that HUQ usually improves upon its base methods, but in some cases, retains the same performance. We suspect that the configurations where HUQ does not outperform the baselines are due to the presence of large covariate shifts between the training and test data. We discuss this in detail in Section 7.
6.2 Overall Comparison
Here, we compare HUQ in selective classification tasks with various other UE techniques, including strong, yet computationally intensive deep ensembles (DE Lakshminarayanan et al., 2017) and SelectiveNet (Geifman and El-Yaniv, 2019) specifically designed for selective classification, but previously tested only in computer vision. Figure 2 presents results for the ELECTRA model and Figure 9 in Appendix D presents results for BERT.
The base epistemic UE methods sometimes cannot outperform even the weak SR baseline or even fall behind it by a large margin. It is especially noticeable for RDE on TWITTER and AMAZON reviews and for DDU on JIGSAW and AMAZON reviews. This effect might appear because the majority of model mistakes arise from ambiguity rather than OOD instances, while these methods are better suitable for OOD detection. On some datasets, it is very hard to overcome the weak SR baseline. For example, on IMPLICITHATE and AMAZON, only DE confidently outperforms SR.
The results for our implementation of SelectiveNet for text classification models and the detailed experimental setup for this method are presented in Appendix F. On all considered datasets, SelectiveNet never outperforms the SR baseline and significantly falls behind it.
Variants of HUQ are usually the best or the second best after DE. For example, HUQ outperforms this strong baseline on PARADETOX, 20 NEWS GROUPS, and SST-5. However, while DE introduces computational overhead of 400%, HUQ requires additionally less than 5% of standard model inference time (see Table 15 in Appendix G).
6.3 Analyses
Hyperparameter for mixing aleatoric and epistemic uncertainty scores in HUQ. When varying the hyperparameter α\alphaα, we change the impact of aleatoric and epistemic uncertainty for the final score. Figure 3 reports the impact of α\alphaα on the TOXIGEN dataset. When α\alphaα is close to 0, the performance of the total score approximates the epistemic uncertainty represented by MD, which is even worse in terms of AUC-RC than the SR baseline. When α\alphaα is close to 1, we use solely the SR score in the mixture of uncertainties, while treating AID, ID, and other instances differently, which results in better performance compared to vanilla SR. The best results are obtained when we select α\alphaα on the validation set. We can see that obtained α=0.5\alpha = 0.5α=0.5 is very close to its optimum on the test set. HUQ-MD in this case outperforms MD by 10.6% and SR by 9.6% in terms of AUC-RC. This again illustrates the importance of mixing different types of uncertainties for selective classification. Similar charts for other considered datasets are presented in Figure 6 in Appendix C and for other hyperparameters in Figures 7 and 8 in Appendix C.
Qualitative analysis. Table 16 in Appendix H presents several instances from various datasets, as well as model predictions and their normalized uncertainty scores. The qualitative analysis reveals that baseline uncertainty scores MD and SR may be high regardless of whether a classification of an instance is correct. For example, we see that four correctly classified instances in PARADETOX are marked with high uncertainty by at least one of the methods. Moreover, MD and SR disagree with each other: MD yields high uncertainty scores for the first two instances, whereas SR produces low uncertainty. For the last two instances, the pattern is reversed. In all of these cases, the MD score is not low enough to consider instances as ID. Therefore, HUQ-MD linearly mixes the SR and MD scores, producing more balanced results with moderately low uncertainty, which is consistent with the fact that classifications are correct.
For the last example from Jigsaw, MD falls below a threshold α\alphaα obtained for this dataset. Consequently, the example is classified as an ID instance, leading to the HUQ-MD score being equal to the SR score for this particular case. Contrary to MD, which yields low uncertainty, high uncertainty of SR and HUQ correctly indicates a prediction error.
For two examples, HUQ-MD contradict the results. Specifically, in the third example of ToxiGen and the second example of Jigsaw, the predictions are accurate, but uncertainty is moderately high. This discrepancy arises from both SR and MD being erroneously high. In such cases, the hybrid method is unable to correct the uncertainty score.
7 Limitations
While HUQ outperforms individual aleatoric and epistemic UE methods for most datasets considered, for some, the effects are negligible. To understand this pattern, we analyze the difference between the training and test sets. We generate latent representations of instances in the datasets using a fine-tuned ELECTRA model and fit a logistic regression model to discriminate between train and test sets using these representations as features. Good performance of the discriminator indicates a covariance shift between the training and test data, while bad performance indicates that instances come from the same distribution.
lt;$SR, DDU
gt;$. As we can see, high F1 scores often correspond to low values of performance gains (the Spearman rank correlation $= 0.8$). This means that HUQ is unlikely to provide improvements to the base methods for the tasks with big covariate shifts. In our analysis, this is due to prediction mistakes primarily arising from OOD instances, which are well-handled by epistemic UE methods.
Visualizing the differences between the datasets using a t-SNE decomposition of the latent representations (see Figure 4), we can see that for IMPLICITHATE and TWITTER, where HUQ does not provide improvements, some regions of the test data are not covered by the training set. For PARADETOX and TOXIGEN, on the other hand, the training dataset completely overlays all regions of the test data, and using HUQ improves AUC-RC on the base methods.
## 8 Conclusion
In this work, we proposed a hybrid uncertainty quantification method for selective text classification. It combines pre-existing methods for aleatoric and epistemic uncertainty, providing scores of total uncertainty. Experimentally, we find that HUQ usually outperforms in terms of RC-AUC other UE methods that aim at quantifying only one type of uncertainty. In real terms, the improved uncertainty estimation offered by our method affords improved identification of erroneous predictions for ambiguous text classification tasks.
Although the HUQ method often provides better results, there are some cases where it is unable to surpass its base methods and performs at a comparable level to them. In our analysis of these examples, we find that this issue arises when there is a substantial covariate shift between the training and test data. In future work, we are planning to analyze other factors that affect the performance of UE methods in selective classification tasks. Our goal is to achieve more consistent and stable improvements over baselines across diverse datasets.
## Acknowledgements
We are very grateful to Zeerak Talat for generously sharing their expertise in toxicity detection, offering valuable suggestions for text edits, and the help with the work in general. We thank anonymous reviewers for their insightful feedback towards improving this paper. The financial support was provided by the Russian Science Foundation, grant 20-71-10135.
## Ethical Considerations
The task of uncertainty estimation is one that is closely tied to the construction of ethical machine learning methods, as it pertains to the identification of potential misclassified instances. For the task of toxic content classification, uncertainty estimation is particularly important due to the speech concerns surrounding toxicity detection. Moreover, toxicity detection has shown disparate performance along gendered and racialized lines, uncertainty estimation provides an avenue for identifying when a model may no longer be applied without further improvement. However, while uncertainty estimation may have potential benefits to the tasks under the umbrella of abusive language detection, approaching misclassifications and uncertainty without an intersectional (Crenshaw, 1991) lens, and without appropriate measures for deep engagements with affected communities may propagate issues of social control, and particularly of enforcing respectability politics of language use. It is therefore important to understand that uncertainty estimation can only provide a partial perspective to the challenges that are faced in abusive language detection. For instance, data that is mislabeled, or labeled such that it propagates stereotypes can exhibit low levels of uncertainty while being undesirable in relation to the goal of equitable machine learning methods for content moderation.
## References
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra. 2015. Weight uncertainty in neural network. In Proceedings of the 32nd International Conference on Machine Learning, volume 37 of Proceedings of Machine Learning Research, pages 1613–1622, Lille, France. PMLR.
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. ELECTRA: pre-training text encoders as discriminators rather than generators. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
Kimberle Crenshaw. 1991. Mapping the margins: Intersectionality, identity politics, and violence against women of color. Stanford Law Review, 43(6):1241–1299.
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017. Automated hate speech detection and the problem of offensive language. Proceedings of the International AAAI Conference on Web and Social Media, 11(1):512–515.
Stefan Depeweg, José Miguel Hernández-Lobato, Finale Doshi-Velez, and Steffen Udluft. 2018. Decomposition of uncertainty in bayesian deep learning for efficient and risk-sensitive learning. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018, volume 80 of Proceedings of Machine Learning Research, pages 1192–1201. PMLR.
Armen Der Kiureghian and Ove Ditlevsen. 2009. Aleatory or epistemic? does it matter? Structural safety, 31(2):105–112.
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Stroudsburg, PA, USA. Association for Computational Linguistics.
Ran El-Yaniv and Yair Wiener. 2010. On the foundations of noise-free selective classification. Journal of Machine Learning Research, 11(53):1605–1641.
Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, and Diyi Yang. 2021. Latent hatred: A benchmark for understanding implicit hate speech. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 345–363, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
Yarin Gal. 2016. Uncertainty in Deep Learning. Ph.D. thesis, University of Cambridge.
Yonatan Geifman and Ran El-Yaniv. 2017. Selective classification for deep neural networks. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NeurIPS 2017, page 4885–4894, Red Hook, NY, USA. Curran Associates Inc.
Yonatan Geifman and Ran El-Yaniv. 2019. SelectiveNet: A deep neural network with an integrated reject option. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, volume 97 of Proceedings of Machine Learning Research, pages 2151–2159. PMLR.
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022. ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3309–3326, Dublin, Ireland. Association for Computational Linguistics.
Jianfeng He, Xuchao Zhang, Shuo Lei, Zhiqian Chen, Fanglan Chen, Abdulaziz Alhamadani, Bei Xiao, and Chang-Tien Lu. 2020. Towards more accurate uncertainty estimation in text classification. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 8362–8372. Association for Computational Linguistics.
Dan Hendrycks and Kevin Gimpel. 2017. A baseline for detecting misclassified and out-of-distribution examples in neural networks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
Nikita Kotelevskii, Aleksandr Artemenkov, Kirill Fedyanin, Fedor Noskov, Alexander Fishkov, Aleksandr Petiushko, Artem Shelmanov, Artem Vazhetsev, and Maxim Panov. 2022. Nonparametric uncertainty quantification for single deterministic neural network. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022.
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. 2017. Simple and scalable predictive uncertainty estimation using deep ensembles. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NeurIPS 2017, page 6405–6416, Red Hook, NY, USA. Curran Associates Inc.
Ken Lang. 1995. Newsweeder: Learning to filter netnews. In Machine Learning, Proceedings of the Twelfth International Conference on Machine Learning, Tahoe City, California, USA, July 9-12, 1995, pages 331–339. Morgan Kaufmann.
Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin. 2018. A simple unified framework for detecting out-of-distribution samples and adversarial attacks. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada, volume 31, pages 7167–7177.
Jeremiah Liu, Zi Lin, Shreyas Padhy, Dustin Tran, Tania Bedrax Weiss, and Balaji Lakshminarayanan. 2020. Simple and principled uncertainty estimation with deterministic deep learning via distance awareness. In Advances in Neural Information Processing Systems, volume 33, pages 7498–7512. Curran Associates, Inc.
Varvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy, David Dale, Irina Krotova, Nikita Semenov, and Alexander Panchenko. 2022. ParaDetox: Detoxification with parallel data. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6804–6818, Dublin, Ireland. Association for Computational Linguistics.
Andrey Malinin and Mark J. F. Gales. 2018. Predictive uncertainty estimation via prior networks. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada, pages 7047–7058.
Julian McAuley and Jure Leskovec. 2013. Hidden factors and hidden topics: Understanding rating dimensions with review text. In Proceedings of the 7th ACM Conference on Recommender Systems, RecSys ’13, page 165–172, New York, NY, USA. Association for Computing Machinery.
Jishnu Mukhoti, Andreas Kirsch, Joost van Amersfoort, Philip HS Torr, and Yarin Gal. 2023. Deep deterministic uncertainty: A new simple baseline. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 24384–24394.
Sarah T. Roberts. 2019. Behind the Screen: Content Moderation in the Shadows of Social Media. Yale University Press.
Peter J Rousseeuw. 1984. Least median of squares regression. Journal of the American statistical association, 79(388):871–880.
Burr Settles. 2009. Active learning literature survey. Computer Sciences Technical Report 1648, University of Wisconsin–Madison.
Artem Shelmanov, Evgenii Tsymbalov, Dmitri Puzyrev, Kirill Fedyanin, Alexander Panchenko, and Maxim Panov. 2021. How certain is your Transformer? In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1833–1840, Online. Association for Computational Linguistics.
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1631–1642, Seattle, Washington, USA. Association for Computational Linguistics.
Sunil Thulasidasan, Gopinath Chennupati, Jeff A. Bilmes, Tanmoy Bhattacharya, and Sarah Michalak. 2019. On mixup training: Improved calibration and predictive uncertainty for deep neural networks. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 13888–13899.
Nanna Thylstrup and Zeerak Waseem. 2020. Detecting ‘dirt’ and ‘toxicity’: Rethinking content moderation as pollution behaviour. SSRN Electronic Journal.
Joost R. van Amersfoort, Lewis Smith, Yee Whye Teh, and Yarin Gal. 2020. Simple and scalable epistemic uncertainty estimation using a single deep deterministic neural network. In International Conference on Machine Learning.
Artem Vazhentsev, Gleb Kuzmin, Artem Shelmanov, Akim Tsvigun, Evgenii Tsymbalov, Kirill Fedyanin, Maxim Panov, Alexander Panchenko, Gleb Gusev, Mikhail Burtsev, Manvel Avetisian, and Leonid Zhukov. 2022. Uncertainty estimation of transformer predictions for misclassification detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8237–8252, Dublin, Ireland. Association for Computational Linguistics.
Zeerak Waseem. 2016. Are you a racist or am I seeing things? annotator influence on hate speech detection on Twitter. In Proceedings of the First Workshop on NLP and Computational Social Science, pages 138–142, Austin, Texas. Association for Computational Linguistics.
Ji Xin, Raphael Tang, Yaoliang Yu, and Jimmy Lin. 2021. The art of abstention: Selective prediction and error regularization for natural language processing. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1040–1051, Online. Association for Computational Linguistics.
KiYoon Yoo, Jangho Kim, Jiho Jang, and Nojun Kwak. 2022. Detection of adversarial examples in text classification: Benchmark and baseline via robust density estimation. In Findings of the Association for Computational Linguistics: ACL 2022, pages 3656–3672, Dublin, Ireland. Association for Computational Linguistics.
Xuchao Zhang, Fanglan Chen, Chang-Tien Lu, and Naren Ramakrishnan. 2019. Mitigating uncertainty in document classification. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3126–3136, Minneapolis, Minnesota. Association for Computational Linguistics.
## A Hyperparameter Values and Hardware Configuration
" data-original-markdown="
Table 4 presents F1 scores for this task aligned with the performance gains of HUQ-DDU in percentages over the best method from the pair
lt;$SR, DDU
gt;$. As we can see, high F1 scores often correspond to low values of performance gains (the Spearman rank correlation $= 0.8$). This means that HUQ is unlikely to provide improvements to the base methods for the tasks with big covariate shifts. In our analysis, this is due to prediction mistakes primarily arising from OOD instances, which are well-handled by epistemic UE methods.
Visualizing the differences between the datasets using a t-SNE decomposition of the latent representations (see Figure 4), we can see that for IMPLICITHATE and TWITTER, where HUQ does not provide improvements, some regions of the test data are not covered by the training set. For PARADETOX and TOXIGEN, on the other hand, the training dataset completely overlays all regions of the test data, and using HUQ improves AUC-RC on the base methods.
## 8 Conclusion
In this work, we proposed a hybrid uncertainty quantification method for selective text classification. It combines pre-existing methods for aleatoric and epistemic uncertainty, providing scores of total uncertainty. Experimentally, we find that HUQ usually outperforms in terms of RC-AUC other UE methods that aim at quantifying only one type of uncertainty. In real terms, the improved uncertainty estimation offered by our method affords improved identification of erroneous predictions for ambiguous text classification tasks.
Although the HUQ method often provides better results, there are some cases where it is unable to surpass its base methods and performs at a comparable level to them. In our analysis of these examples, we find that this issue arises when there is a substantial covariate shift between the training and test data. In future work, we are planning to analyze other factors that affect the performance of UE methods in selective classification tasks. Our goal is to achieve more consistent and stable improvements over baselines across diverse datasets.
## Acknowledgements
We are very grateful to Zeerak Talat for generously sharing their expertise in toxicity detection, offering valuable suggestions for text edits, and the help with the work in general. We thank anonymous reviewers for their insightful feedback towards improving this paper. The financial support was provided by the Russian Science Foundation, grant 20-71-10135.
## Ethical Considerations
The task of uncertainty estimation is one that is closely tied to the construction of ethical machine learning methods, as it pertains to the identification of potential misclassified instances. For the task of toxic content classification, uncertainty estimation is particularly important due to the speech concerns surrounding toxicity detection. Moreover, toxicity detection has shown disparate performance along gendered and racialized lines, uncertainty estimation provides an avenue for identifying when a model may no longer be applied without further improvement. However, while uncertainty estimation may have potential benefits to the tasks under the umbrella of abusive language detection, approaching misclassifications and uncertainty without an intersectional (Crenshaw, 1991) lens, and without appropriate measures for deep engagements with affected communities may propagate issues of social control, and particularly of enforcing respectability politics of language use. It is therefore important to understand that uncertainty estimation can only provide a partial perspective to the challenges that are faced in abusive language detection. For instance, data that is mislabeled, or labeled such that it propagates stereotypes can exhibit low levels of uncertainty while being undesirable in relation to the goal of equitable machine learning methods for content moderation.
## References
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra. 2015. Weight uncertainty in neural network. In Proceedings of the 32nd International Conference on Machine Learning, volume 37 of Proceedings of Machine Learning Research, pages 1613–1622, Lille, France. PMLR.
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. ELECTRA: pre-training text encoders as discriminators rather than generators. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
Kimberle Crenshaw. 1991. Mapping the margins: Intersectionality, identity politics, and violence against women of color. Stanford Law Review, 43(6):1241–1299.
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017. Automated hate speech detection and the problem of offensive language. Proceedings of the International AAAI Conference on Web and Social Media, 11(1):512–515.
Stefan Depeweg, José Miguel Hernández-Lobato, Finale Doshi-Velez, and Steffen Udluft. 2018. Decomposition of uncertainty in bayesian deep learning for efficient and risk-sensitive learning. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018, volume 80 of Proceedings of Machine Learning Research, pages 1192–1201. PMLR.
Armen Der Kiureghian and Ove Ditlevsen. 2009. Aleatory or epistemic? does it matter? Structural safety, 31(2):105–112.
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Stroudsburg, PA, USA. Association for Computational Linguistics.
Ran El-Yaniv and Yair Wiener. 2010. On the foundations of noise-free selective classification. Journal of Machine Learning Research, 11(53):1605–1641.
Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, and Diyi Yang. 2021. Latent hatred: A benchmark for understanding implicit hate speech. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 345–363, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
Yarin Gal. 2016. Uncertainty in Deep Learning. Ph.D. thesis, University of Cambridge.
Yonatan Geifman and Ran El-Yaniv. 2017. Selective classification for deep neural networks. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NeurIPS 2017, page 4885–4894, Red Hook, NY, USA. Curran Associates Inc.
Yonatan Geifman and Ran El-Yaniv. 2019. SelectiveNet: A deep neural network with an integrated reject option. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, volume 97 of Proceedings of Machine Learning Research, pages 2151–2159. PMLR.
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022. ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3309–3326, Dublin, Ireland. Association for Computational Linguistics.
Jianfeng He, Xuchao Zhang, Shuo Lei, Zhiqian Chen, Fanglan Chen, Abdulaziz Alhamadani, Bei Xiao, and Chang-Tien Lu. 2020. Towards more accurate uncertainty estimation in text classification. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 8362–8372. Association for Computational Linguistics.
Dan Hendrycks and Kevin Gimpel. 2017. A baseline for detecting misclassified and out-of-distribution examples in neural networks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
Nikita Kotelevskii, Aleksandr Artemenkov, Kirill Fedyanin, Fedor Noskov, Alexander Fishkov, Aleksandr Petiushko, Artem Shelmanov, Artem Vazhetsev, and Maxim Panov. 2022. Nonparametric uncertainty quantification for single deterministic neural network. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022.
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. 2017. Simple and scalable predictive uncertainty estimation using deep ensembles. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NeurIPS 2017, page 6405–6416, Red Hook, NY, USA. Curran Associates Inc.
Ken Lang. 1995. Newsweeder: Learning to filter netnews. In Machine Learning, Proceedings of the Twelfth International Conference on Machine Learning, Tahoe City, California, USA, July 9-12, 1995, pages 331–339. Morgan Kaufmann.
Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin. 2018. A simple unified framework for detecting out-of-distribution samples and adversarial attacks. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada, volume 31, pages 7167–7177.
Jeremiah Liu, Zi Lin, Shreyas Padhy, Dustin Tran, Tania Bedrax Weiss, and Balaji Lakshminarayanan. 2020. Simple and principled uncertainty estimation with deterministic deep learning via distance awareness. In Advances in Neural Information Processing Systems, volume 33, pages 7498–7512. Curran Associates, Inc.
Varvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy, David Dale, Irina Krotova, Nikita Semenov, and Alexander Panchenko. 2022. ParaDetox: Detoxification with parallel data. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6804–6818, Dublin, Ireland. Association for Computational Linguistics.
Andrey Malinin and Mark J. F. Gales. 2018. Predictive uncertainty estimation via prior networks. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada, pages 7047–7058.
Julian McAuley and Jure Leskovec. 2013. Hidden factors and hidden topics: Understanding rating dimensions with review text. In Proceedings of the 7th ACM Conference on Recommender Systems, RecSys ’13, page 165–172, New York, NY, USA. Association for Computing Machinery.
Jishnu Mukhoti, Andreas Kirsch, Joost van Amersfoort, Philip HS Torr, and Yarin Gal. 2023. Deep deterministic uncertainty: A new simple baseline. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 24384–24394.
Sarah T. Roberts. 2019. Behind the Screen: Content Moderation in the Shadows of Social Media. Yale University Press.
Peter J Rousseeuw. 1984. Least median of squares regression. Journal of the American statistical association, 79(388):871–880.
Burr Settles. 2009. Active learning literature survey. Computer Sciences Technical Report 1648, University of Wisconsin–Madison.
Artem Shelmanov, Evgenii Tsymbalov, Dmitri Puzyrev, Kirill Fedyanin, Alexander Panchenko, and Maxim Panov. 2021. How certain is your Transformer? In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1833–1840, Online. Association for Computational Linguistics.
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1631–1642, Seattle, Washington, USA. Association for Computational Linguistics.
Sunil Thulasidasan, Gopinath Chennupati, Jeff A. Bilmes, Tanmoy Bhattacharya, and Sarah Michalak. 2019. On mixup training: Improved calibration and predictive uncertainty for deep neural networks. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 13888–13899.
Nanna Thylstrup and Zeerak Waseem. 2020. Detecting ‘dirt’ and ‘toxicity’: Rethinking content moderation as pollution behaviour. SSRN Electronic Journal.
Joost R. van Amersfoort, Lewis Smith, Yee Whye Teh, and Yarin Gal. 2020. Simple and scalable epistemic uncertainty estimation using a single deep deterministic neural network. In International Conference on Machine Learning.
Artem Vazhentsev, Gleb Kuzmin, Artem Shelmanov, Akim Tsvigun, Evgenii Tsymbalov, Kirill Fedyanin, Maxim Panov, Alexander Panchenko, Gleb Gusev, Mikhail Burtsev, Manvel Avetisian, and Leonid Zhukov. 2022. Uncertainty estimation of transformer predictions for misclassification detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8237–8252, Dublin, Ireland. Association for Computational Linguistics.
Zeerak Waseem. 2016. Are you a racist or am I seeing things? annotator influence on hate speech detection on Twitter. In Proceedings of the First Workshop on NLP and Computational Social Science, pages 138–142, Austin, Texas. Association for Computational Linguistics.
Ji Xin, Raphael Tang, Yaoliang Yu, and Jimmy Lin. 2021. The art of abstention: Selective prediction and error regularization for natural language processing. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1040–1051, Online. Association for Computational Linguistics.
KiYoon Yoo, Jangho Kim, Jiho Jang, and Nojun Kwak. 2022. Detection of adversarial examples in text classification: Benchmark and baseline via robust density estimation. In Findings of the Association for Computational Linguistics: ACL 2022, pages 3656–3672, Dublin, Ireland. Association for Computational Linguistics.
Xuchao Zhang, Fanglan Chen, Chang-Tien Lu, and Naren Ramakrishnan. 2019. Mitigating uncertainty in document classification. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3126–3136, Minneapolis, Minnesota. Association for Computational Linguistics.
## A Hyperparameter Values and Hardware Configuration
" data-source-offset="37938" class="markdown-segment">
Table 4 presents F1 scores for this task aligned with the performance gains of HUQ-DDU in percentages over the best method from the pair <<<SR, DDU>>>. As we can see, high F1 scores often correspond to low values of performance gains (the Spearman rank correlation =0.8= 0.8=0.8). This means that HUQ is unlikely to provide improvements to the base methods for the tasks with big covariate shifts. In our analysis, this is due to prediction mistakes primarily arising from OOD instances, which are well-handled by epistemic UE methods.
Visualizing the differences between the datasets using a t-SNE decomposition of the latent representations (see Figure 4), we can see that for IMPLICITHATE and TWITTER, where HUQ does not provide improvements, some regions of the test data are not covered by the training set. For PARADETOX and TOXIGEN, on the other hand, the training dataset completely overlays all regions of the test data, and using HUQ improves AUC-RC on the base methods.
8 Conclusion
In this work, we proposed a hybrid uncertainty quantification method for selective text classification. It combines pre-existing methods for aleatoric and epistemic uncertainty, providing scores of total uncertainty. Experimentally, we find that HUQ usually outperforms in terms of RC-AUC other UE methods that aim at quantifying only one type of uncertainty. In real terms, the improved uncertainty estimation offered by our method affords improved identification of erroneous predictions for ambiguous text classification tasks.
Although the HUQ method often provides better results, there are some cases where it is unable to surpass its base methods and performs at a comparable level to them. In our analysis of these examples, we find that this issue arises when there is a substantial covariate shift between the training and test data. In future work, we are planning to analyze other factors that affect the performance of UE methods in selective classification tasks. Our goal is to achieve more consistent and stable improvements over baselines across diverse datasets.
Acknowledgements
We are very grateful to Zeerak Talat for generously sharing their expertise in toxicity detection, offering valuable suggestions for text edits, and the help with the work in general. We thank anonymous reviewers for their insightful feedback towards improving this paper. The financial support was provided by the Russian Science Foundation, grant 20-71-10135.
Ethical Considerations
The task of uncertainty estimation is one that is closely tied to the construction of ethical machine learning methods, as it pertains to the identification of potential misclassified instances. For the task of toxic content classification, uncertainty estimation is particularly important due to the speech concerns surrounding toxicity detection. Moreover, toxicity detection has shown disparate performance along gendered and racialized lines, uncertainty estimation provides an avenue for identifying when a model may no longer be applied without further improvement. However, while uncertainty estimation may have potential benefits to the tasks under the umbrella of abusive language detection, approaching misclassifications and uncertainty without an intersectional (Crenshaw, 1991) lens, and without appropriate measures for deep engagements with affected communities may propagate issues of social control, and particularly of enforcing respectability politics of language use. It is therefore important to understand that uncertainty estimation can only provide a partial perspective to the challenges that are faced in abusive language detection. For instance, data that is mislabeled, or labeled such that it propagates stereotypes can exhibit low levels of uncertainty while being undesirable in relation to the goal of equitable machine learning methods for content moderation.
References
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra. 2015. Weight uncertainty in neural network. In Proceedings of the 32nd International Conference on Machine Learning, volume 37 of Proceedings of Machine Learning Research, pages 1613–1622, Lille, France. PMLR.
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020. ELECTRA: pre-training text encoders as discriminators rather than generators. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
Kimberle Crenshaw. 1991. Mapping the margins: Intersectionality, identity politics, and violence against women of color. Stanford Law Review, 43(6):1241–1299.
Thomas Davidson, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017. Automated hate speech detection and the problem of offensive language. Proceedings of the International AAAI Conference on Web and Social Media, 11(1):512–515.
Stefan Depeweg, José Miguel Hernández-Lobato, Finale Doshi-Velez, and Steffen Udluft. 2018. Decomposition of uncertainty in bayesian deep learning for efficient and risk-sensitive learning. In Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018, volume 80 of Proceedings of Machine Learning Research, pages 1192–1201. PMLR.
Armen Der Kiureghian and Ove Ditlevsen. 2009. Aleatory or epistemic? does it matter? Structural safety, 31(2):105–112.
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Stroudsburg, PA, USA. Association for Computational Linguistics.
Ran El-Yaniv and Yair Wiener. 2010. On the foundations of noise-free selective classification. Journal of Machine Learning Research, 11(53):1605–1641.
Mai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi, Jordyn Seybolt, Munmun De Choudhury, and Diyi Yang. 2021. Latent hatred: A benchmark for understanding implicit hate speech. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 345–363, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
Yarin Gal. 2016. Uncertainty in Deep Learning. Ph.D. thesis, University of Cambridge.
Yonatan Geifman and Ran El-Yaniv. 2017. Selective classification for deep neural networks. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NeurIPS 2017, page 4885–4894, Red Hook, NY, USA. Curran Associates Inc.
Yonatan Geifman and Ran El-Yaniv. 2019. SelectiveNet: A deep neural network with an integrated reject option. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, volume 97 of Proceedings of Machine Learning Research, pages 2151–2159. PMLR.
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar. 2022. ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3309–3326, Dublin, Ireland. Association for Computational Linguistics.
Jianfeng He, Xuchao Zhang, Shuo Lei, Zhiqian Chen, Fanglan Chen, Abdulaziz Alhamadani, Bei Xiao, and Chang-Tien Lu. 2020. Towards more accurate uncertainty estimation in text classification. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 8362–8372. Association for Computational Linguistics.
Dan Hendrycks and Kevin Gimpel. 2017. A baseline for detecting misclassified and out-of-distribution examples in neural networks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
Nikita Kotelevskii, Aleksandr Artemenkov, Kirill Fedyanin, Fedor Noskov, Alexander Fishkov, Aleksandr Petiushko, Artem Shelmanov, Artem Vazhetsev, and Maxim Panov. 2022. Nonparametric uncertainty quantification for single deterministic neural network. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022.
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. 2017. Simple and scalable predictive uncertainty estimation using deep ensembles. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NeurIPS 2017, page 6405–6416, Red Hook, NY, USA. Curran Associates Inc.
Ken Lang. 1995. Newsweeder: Learning to filter netnews. In Machine Learning, Proceedings of the Twelfth International Conference on Machine Learning, Tahoe City, California, USA, July 9-12, 1995, pages 331–339. Morgan Kaufmann.
Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin. 2018. A simple unified framework for detecting out-of-distribution samples and adversarial attacks. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada, volume 31, pages 7167–7177.
Jeremiah Liu, Zi Lin, Shreyas Padhy, Dustin Tran, Tania Bedrax Weiss, and Balaji Lakshminarayanan. 2020. Simple and principled uncertainty estimation with deterministic deep learning via distance awareness. In Advances in Neural Information Processing Systems, volume 33, pages 7498–7512. Curran Associates, Inc.
Varvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy, David Dale, Irina Krotova, Nikita Semenov, and Alexander Panchenko. 2022. ParaDetox: Detoxification with parallel data. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6804–6818, Dublin, Ireland. Association for Computational Linguistics.
Andrey Malinin and Mark J. F. Gales. 2018. Predictive uncertainty estimation via prior networks. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada, pages 7047–7058.
Julian McAuley and Jure Leskovec. 2013. Hidden factors and hidden topics: Understanding rating dimensions with review text. In Proceedings of the 7th ACM Conference on Recommender Systems, RecSys ’13, page 165–172, New York, NY, USA. Association for Computing Machinery.
Jishnu Mukhoti, Andreas Kirsch, Joost van Amersfoort, Philip HS Torr, and Yarin Gal. 2023. Deep deterministic uncertainty: A new simple baseline. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 24384–24394.
Sarah T. Roberts. 2019. Behind the Screen: Content Moderation in the Shadows of Social Media. Yale University Press.
Peter J Rousseeuw. 1984. Least median of squares regression. Journal of the American statistical association, 79(388):871–880.
Burr Settles. 2009. Active learning literature survey. Computer Sciences Technical Report 1648, University of Wisconsin–Madison.
Artem Shelmanov, Evgenii Tsymbalov, Dmitri Puzyrev, Kirill Fedyanin, Alexander Panchenko, and Maxim Panov. 2021. How certain is your Transformer? In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1833–1840, Online. Association for Computational Linguistics.
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1631–1642, Seattle, Washington, USA. Association for Computational Linguistics.
Sunil Thulasidasan, Gopinath Chennupati, Jeff A. Bilmes, Tanmoy Bhattacharya, and Sarah Michalak. 2019. On mixup training: Improved calibration and predictive uncertainty for deep neural networks. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, pages 13888–13899.
Nanna Thylstrup and Zeerak Waseem. 2020. Detecting ‘dirt’ and ‘toxicity’: Rethinking content moderation as pollution behaviour. SSRN Electronic Journal.
Joost R. van Amersfoort, Lewis Smith, Yee Whye Teh, and Yarin Gal. 2020. Simple and scalable epistemic uncertainty estimation using a single deep deterministic neural network. In International Conference on Machine Learning.
Artem Vazhentsev, Gleb Kuzmin, Artem Shelmanov, Akim Tsvigun, Evgenii Tsymbalov, Kirill Fedyanin, Maxim Panov, Alexander Panchenko, Gleb Gusev, Mikhail Burtsev, Manvel Avetisian, and Leonid Zhukov. 2022. Uncertainty estimation of transformer predictions for misclassification detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8237–8252, Dublin, Ireland. Association for Computational Linguistics.
Zeerak Waseem. 2016. Are you a racist or am I seeing things? annotator influence on hate speech detection on Twitter. In Proceedings of the First Workshop on NLP and Computational Social Science, pages 138–142, Austin, Texas. Association for Computational Linguistics.
Ji Xin, Raphael Tang, Yaoliang Yu, and Jimmy Lin. 2021. The art of abstention: Selective prediction and error regularization for natural language processing. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1040–1051, Online. Association for Computational Linguistics.
KiYoon Yoo, Jangho Kim, Jiho Jang, and Nojun Kwak. 2022. Detection of adversarial examples in text classification: Benchmark and baseline via robust density estimation. In Findings of the Association for Computational Linguistics: ACL 2022, pages 3656–3672, Dublin, Ireland. Association for Computational Linguistics.
Xuchao Zhang, Fanglan Chen, Chang-Tien Lu, and Naren Ramakrishnan. 2019. Mitigating uncertainty in document classification. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3126–3136, Minneapolis, Minnesota. Association for Computational Linguistics.
A Hyperparameter Values and Hardware Configuration