Built independently by an author, for readers. Read the story and support ChapterPal

keyword

text classification

Text classification is a natural language processing and machine learning task that involves automatically assigning one or more predefined categories or labels to unstructured text documents, sentences, or queries based on their content. The process converts raw text into numerical representations, which can range from traditional statistical features and word vectors to dense representations derived from deep neural networks and pre-trained transformer language models. Algorithms then analyze these features to identify patterns and predict the appropriate classes, supporting diverse real-world applications such as sentiment analysis, spam detection, news topic categorization, user intent recognition, and content moderation.

43 items

Bag-of-Words vs. Graph vs. Sequence in Text Classification: Questioning the Necessity of Text-Graphs and the Surprising Strength of a Wide MLP

Bag-of-Words vs. Graph vs. Sequence in Text Classification: Questioning the Necessity of Text-Graphs and the Surprising Strength of a Wide MLP

Lukas Galke, Ansgar Scherp

OrganizationsMax Planck Institute for PsycholinguisticsUlm UniversityUniversity of Kiel

Why you should read this

Demonstrates that a simple, wide bag-of-words multi-layer perceptron outperforms complex graph neural networks like TextGCN in inductive text classification while offering significantly faster training and inference than Transformer models on long sequences.

Graph neural networks have triggered a resurgence of graph-based text classification methods, defining today’s state of the art. We show that a wide multi-layer perceptron (MLP) using a Bag-of-Words (BoW) outperforms the recent graph-based models TextGCN and HeteGCN in an inductive text classification setting and is comparable with HyperGAT. Moreover, we fine-tune a sequence-based BERT and a lightweight DistilBERT model, which both outperform all state-of-the-art models. These results question the importance of synthetic graphs used in modern text classifiers. In terms of efficiency, DistilBERT is still twice as large as our BoW-based wide MLP, while graph-based models like TextGCN require setting up an O(N²) graph, where N is the vocabulary plus corpus size. Finally, since Transformers need to compute O(L²) attention weights with sequence length L, the MLP models show higher training and inference speeds on datasets with long sequences.

Added

2026-10-03

DoCoGen: Domain Counterfactual Generation for Low Resource Domain Adaptation

DoCoGen: Domain Counterfactual Generation for Low Resource Domain Adaptation

Nitay Calderon, Eyal Ben-David, Amir Feder, Roi Reichart

OrganizationsTechnion – Israel Institute of Technology

Why you should read this

Proposes DoCoGen, an unsupervised controllable generation framework that transforms multi-sentence texts across domains while preserving their task labels, enabling effective data augmentation for low-resource domain adaptation without requiring parallel text or target task annotations.

Natural language processing (NLP) algorithms have become very successful, but they still struggle when applied to out-of-distribution examples. In this paper we propose a controllable generation approach in order to deal with this domain adaptation (DA) challenge. Given an input text example, our DoCoGen algorithm generates a domain-counterfactual textual example (D-CON) – that is similar to the original in all aspects, including the task label, but its domain is changed to a desired one. Importantly, DoCoGen is trained using only unlabeled examples from multiple domains – no NLP task labels or parallel pairs of textual examples and their domain-counterfactuals are required. We show that DoCoGen can generate coherent counterfactuals consisting of multiple sentences. We use the D-CONs generated by DoCoGen to augment a sentiment classifier and a multi-label intent classifier in 20 and 78 DA setups, respectively, where source-domain labeled data is scarce. Our model outperforms strong baselines and improves the accuracy of a state-of-the-art unsupervised DA algorithm.

Added

2026-10-03

Explore Spurious Correlations at the Concept Level in Language Models for Text Classification

Explore Spurious Correlations at the Concept Level in Language Models for Text Classification

Yuhang Zhou, Paiheng Xu, Xiaoyu Liu, Bang An, Wei Ai, Furong Huang

OrganizationsUniversity of Maryland

Why you should read this

Reveals how language models learn shortcut predictions from broader concept-level biases in both fine-tuning and in-context learning, and presents a counterfactual data rebalancing method to eliminate these errors while preserving classification accuracy.

Language models (LMs) have achieved notable success in numerous NLP tasks, employing both fine-tuning and in-context learning (ICL) methods. While language models demonstrate exceptional performance, they face robustness challenges due to spurious correlations arising from imbalanced label distributions in training data or ICL exemplars. Previous research has primarily concentrated on word, phrase, and syntax features, neglecting the concept level, often due to the absence of concept labels and difficulty in identifying conceptual content in input texts. This paper introduces two main contributions. First, we employ ChatGPT to assign concept labels to texts, assessing concept bias in models during fine-tuning or ICL on test data. We find that LMs, when encountering spurious correlations between a concept and a label in training or prompts, resort to shortcuts for predictions. Second, we introduce a data re-balancing technique that incorporates ChatGPT-generated counterfactual data, thereby balancing label distribution and mitigating spurious correlations. Our method’s efficacy, surpassing traditional token removal approaches, is validated through extensive testing.

Added

2026-10-03

"Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification

"Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification

Jasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm, Katja Filippova

Why you should read this

Establishes a rigorous evaluation protocol using synthetic shortcut injection to benchmark the faithfulness of input salience methods, revealing that popular explanation techniques often fail to detect even simple lexical patterns used by text classifiers.

A common approach for explaining predictions made by neural networks is to identify salient features using gradients or attention. In natural language processing specifically, there have been several approaches proposed to compute gradient-based token importance scores; yet their faithfulness remains unclear: e.g., do they agree where important tokens occur? different formulations may yield completely different explanations, making comparisons across papers ambiguous. Existing evaluation practice often relies solely on intuition—both when designing new explanation techniques and reporting results.We propose experiments called “shortcut” tasks specially crafted such that we know which spurious patterns could exist in data and how easily models would pick them up instead of meaningful signals.Empirically, neither commonly used variants of input salience methods nor attention pass our tests consistently.This calls into question prior work based entirely on automatic metrics.Finally, we suggest simple practical recommendations relating design choices in evaluating model behavior to obtainmore faithful attributions.The code base supporting all analyses reported here hasbeen released alongwith an interactive online tool accompanyingthis paper.Both links accessibleat https://github.com/google-research/ \allowbreak google-research/tree/master/saliency\_faith fulness .Thus,in summary,the contributionsofourpaperare( )a setoffive shortcuttasksincontrolledsettingsofdifferentcomplexities;(ii) asystematicstudyofgradientbasedsaliencymethodsacrossthese five settings,w.r.t.multipletextclassifiersincludingstateoftheartmodelsfine tunedonfourlanguage datasets,and(iii)a concrete proposaltoaugmentexistingbenchmarksingle-taskmetricsbyadditionalmulti taskediagnosticsdrivenbyeitheroftheshortcutdatasetsintroducedorrealisticdomainspecificpatternsaswe exemplifyfortoxicitydetection.Wenotethatexplanationsamplersemustbecausally relatedtothedecisionprocess,butalsomusto beclearly communicatedtothestakeholder,i.e.acontentprovideroranend-user-inthatsenseouremphasisisonanalyzingthemethodratherthanjustperformance.Hence,intheabsenceofaformalframeworkforevaluatingex planation qualityweproposeaprotocolthatisolatesbehavioraldifferenceswhile encouragingtheadoptionofcommonscoresfordistinctclasses ofexplanation methods

Added

2026-10-02

Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations

Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations

Zhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ming Yin

OrganizationsPurdue UniversityWashington University in St. Louis

Why you should read this

Demonstrates through empirical evaluation across ten datasets that the effectiveness of LLM-generated synthetic training data for text classifiers degrades significantly as task and instance subjectivity increase.

The collection and curation of high-quality training data is crucial for developing text classification models with superior performance, but it is often associated with significant costs and time investment. Researchers have recently explored using large language models (LLMs) to generate synthetic datasets as an alternative approach. However, the effectiveness of the LLM-generated synthetic data in supporting model training is inconsistent across different classification tasks. To better understand factors that moderate the effectiveness of the LLM-generated synthetic data, in this study, we look into how the performance of models trained on these synthetic data may vary with the subjectivity of classification. Our results indicate that subjectivity, at both the task level and instance level, is negatively associated with the performance of the model trained on synthetic data. We conclude by discussing the implications of our work on the potential and limitations of leveraging LLM for synthetic data generation¹.

Added

2026-10-01

NoisyTune: A Little Noise Can Help You Finetune Pretrained Language Models Better

NoisyTune: A Little Noise Can Help You Finetune Pretrained Language Models Better

Chuhan Wu, Fangzhao Wu, Tao Qi, Yongfeng Huang

OrganizationsMicrosoftTsinghua University

Why you should read this

Proposes NoisyTune, a lightweight and easily implementable method that improves downstream fine-tuning by adding matrix-wise uniform noise scaled by parameter standard deviations to prevent pretrained language models from overfitting their pretraining tasks.

Effectively finetuning pretrained language models (PLMs) is critical for their success in downstream tasks. However, PLMs may have risks in overfitting the pretraining tasks and data, which usually have gap with the target downstream tasks. Such gap may be difficult for existing PLM finetuning methods to overcome and lead to suboptimal performance. In this paper, we propose a very simple yet effective method named NoisyTune to help better finetune PLMs on downstream tasks by adding some noise to the parameters of PLMs before finetuning. More specifically, we propose a matrix-wise perturbing method which adds different uniform noises to different parameter matrices based on their standard deviations. In this way, the varied characteristics of different types of parameters in PLMs can be considered. Extensive experiments on both GLUE English benchmark and XTREME multilingual benchmark show NoisyTune can consistently empower the finetuning of different PLMs on different downstream tasks.

Added

2026-09-26

TextHoaxer: Budgeted Hard-Label Adversarial Attacks on Text

TextHoaxer: Budgeted Hard-Label Adversarial Attacks on Text

Muchao Ye, Chenglin Miao, Ting Wang, Fenglong Ma

OrganizationsPennsylvania State UniversityUniversity of Georgia

Why you should read this

Proposes TextHoaxer, a gradient-based framework that formulates hard-label text adversarial attacks in continuous word embedding space to generate high-similarity adversarial examples under strict query budgets without relying on query-expensive genetic algorithms.

This paper focuses on a newly challenging setting in hard-label adversarial attacks on text data by taking the budget information into account. Although existing approaches can successfully generate adversarial examples in the hard-label setting, they follow an ideal assumption that the victim model does not restrict the number of queries. However, in real-world applications the query budget is usually tight or limited. Moreover, existing hard-label adversarial attack techniques use the genetic algorithm to optimize discrete text data by maintaining a number of adversarial candidates during optimization, which can lead to the problem of generating low-quality adversarial examples in the tight-budget setting. To solve this problem, in this paper, we propose a new method named TextHoaxer by formulating the budgeted hard-label adversarial attack task on text data as a gradient-based optimization problem of perturbation matrix in the continuous word embedding space. Compared with the genetic algorithm-based optimization, our solution only uses a single initialized adversarial example as the adversarial candidate for optimization, which significantly reduces the number of queries. The optimization is guided by a new objective function consisting of three terms, i.e., semantic similarity term, pair-wise perturbation constraint, and sparsity constraint. Semantic similarity term and pair-wise perturbation constraint can ensure the high semantic similarity of adversarial examples from both comprehensive text-level and individual word-level, while the sparsity constraint explicitly restricts the number of perturbed words, which is also helpful for enhancing the quality of generated text. We conduct extensive experiments on eight text datasets against three representative natural language models, and experimental results show that TextHoaxer can generate high-quality adversarial examples with higher semantic similarity and lower perturbation rate under the tight-budget setting.

Added

2026-09-26

Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions

Increasing Diversity While Maintaining Accuracy: Text Data Generation with Large Language Models and Human Interventions

John Joon Young Chung, Ece Kamar, Saleema Amershi

OrganizationsMicrosoftUniversity of Michigan

Why you should read this

Demonstrates how combining large language model diversification techniques with targeted human label replacement allows smaller downstream classifiers to outperform few-shot large language models by balancing synthetic text variety and annotation accuracy.

Large language models (LLMs) can be used to generate text data for training and evaluating other models. However, creating high-quality datasets with LLMs can be challenging. In this work, we explore human-AI partnerships to facilitate high diversity and accuracy in LLM-based text data generation. We first examine two approaches to diversify text generation: 1) logit suppression, which minimizes the generation of languages that have already been frequently generated, and 2) temperature sampling, which flattens the token sampling probability. We found that diversification approaches can increase data diversity but often at the cost of data accuracy (i.e., text and labels being appropriate for the target domain). To address this issue, we examined two human interventions, 1) label replacement (LR), correcting misaligned labels, and 2) out-of-scope filtering (OOSF), removing instances that are out of the user’s domain of interest or to which no considered label applies. With oracle studies, we found that LR increases the absolute accuracy of models trained with diversified datasets by 14.4%. Moreover, we found that some models trained with data generated with LR interventions outperformed LLM-based few-shot classification. In contrast, OOSF was not effective in increasing model accuracy, implying the need for future work in human-in-the-loop text data generation.

Added

2026-09-26

A Sensitivity Analysis of (and Practitioners’ Guide to) Convolutional Neural Networks for Sentence Classification

A Sensitivity Analysis of (and Practitioners’ Guide to) Convolutional Neural Networks for Sentence Classification

Ye Zhang, Byron C. Wallace

OrganizationsUniversity of Texas at Austin

Why you should read this

Establishes practical guidelines for configuring convolutional neural networks in sentence classification by isolating which hyperparameter choices critically influence model accuracy and which can be safely neglected.

Convolutional Neural Networks (CNNs) have recently achieved remarkably strong performance on the practically important task of sentence classification (kim 2014, kalchbrenner 2014, johnson 2014). However, these models require practitioners to specify an exact model architecture and set accompanying hyperparameters, including the filter region size, regularization parameters, and so on. It is currently unknown how sensitive model performance is to changes in these configurations for the task of sentence classification. We thus conduct a sensitivity analysis of one-layer CNNs to explore the effect of architecture components on model performance; our aim is to distinguish between important and comparatively inconsequential design decisions for sentence classification. We focus on one-layer CNNs (to the exclusion of more complex models) due to their comparative simplicity and strong empirical performance, which makes it a modern standard baseline method akin to Support Vector Machine (SVMs) and logistic regression. We derive practical advice from our extensive empirical results for those interested in getting the most out of CNNs for sentence classification in real world settings.

Added

2026-09-25

Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment

Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment

Di Jin, Zhijing Jin, Joey Tianyi Zhou, Peter Szolovits

OrganizationsAgency for Science, Technology and ResearchMassachusetts Institute of TechnologyUniversity of Hong Kong

Why you should read this

Proposes TextFooler, a simple and computationally efficient attack framework that systematically deceives pre-trained models like BERT by generating natural, semantics-preserving adversarial examples for classification and textual entailment.

Machine learning algorithms are often vulnerable to adversarial examples that have imperceptible alterations from the original counterparts but can fool the state-of-the-art models. It is helpful to evaluate or even improve the robustness of these models by exposing the maliciously crafted adversarial examples. In this paper, we present TextFooler, a simple but strong baseline to generate natural adversarial text. By applying it to two fundamental natural language tasks, text classification and textual entailment, we successfully attacked three target models, including the powerful pre-trained BERT, and the widely used convolutional and recurrent neural networks. We demonstrate the advantages of this framework in three ways: (1) effective---it outperforms state-of-the-art attacks in terms of success rate and perturbation rate, (2) utility-preserving---it preserves semantic content and grammaticality, and remains correctly classified by humans, and (3) efficient---it generates adversarial text with computational complexity linear to the text length. *The code, pre-trained target models, and test examples are available at this https URL.

Added

2026-09-25

Text Classification using String Kernels

Text Classification using String Kernels

H. Lodhi, C. Saunders, J. Shawe-Taylor, N. Cristianini, Christopher J. C. H. Watkins

OrganizationsRoyal Holloway, University of London

Why you should read this

Proposes a string subsequence kernel for text categorization that captures non-contiguous character sequences via dynamic programming and outperforms traditional bag-of-words representations on benchmark corpora.

We propose a novel approach for categorizing text documents based on the use of a special kernel. The kernel is an inner product in the feature space generated by all subsequences of length k. A subsequence is any ordered sequence of k characters occurring in the text though not necessarily contiguously. The subsequences are weighted by an exponentially decaying factor of their full length in the text, hence emphasising those occurrences that are close to contiguous. A direct computation of this feature vector would involve a prohibitive amount of computation even for modest values of k, since the dimension of the feature space grows exponentially with k. The paper describes how despite this fact the inner product can be efficiently evaluated by a dynamic programming technique. Experimental comparisons of the performance of the kernel compared with a standard word feature space kernel (Joachims, 1998) show positive results on modestly sized datasets. The case of contiguous subsequences is also considered for comparison with the subsequences kernel with different decay factors. For larger documents and datasets the paper introduces an approximation technique that is shown to deliver good approximations efficiently for large datasets.

Added

2026-09-25

Measuring praise and criticism: Inference of semantic orientation from association

Measuring praise and criticism: Inference of semantic orientation from association

Peter D. Turney, Michael L. Littman

OrganizationsNational Research Council CanadaRutgers University

Why you should read this

Proposes a method for automatically determining the positive or negative semantic orientation of words across diverse parts of speech by measuring their statistical association with paradigm seed words using pointwise mutual information and latent semantic analysis.

The evaluative character of a word is called its semantic orientation. Positive semantic orientation indicates praise (e.g., "honest", "intrepid") and negative semantic orientation indicates criticism (e.g., "disturbing", "superfluous"). Semantic orientation varies in both direction (positive or negative) and degree (mild to strong). An automated system for measuring semantic orientation would have application in text classification, text filtering, tracking opinions in online discussions, analysis of survey responses, and automated chat systems (chatbots). This paper introduces a method for inferring the semantic orientation of a word from its statistical association with a set of positive and negative paradigm words. Two instances of this approach are evaluated, based on two different statistical measures of word association: pointwise mutual information (PMI) and latent semantic analysis (LSA). The method is experimentally tested with 3,596 words (including adjectives, adverbs, nouns, and verbs) that have been manually labeled positive (1,614 words) and negative (1,982 words). The method attains an accuracy of 82.8% on the full test set, but the accuracy rises above 95% when the algorithm is allowed to abstain from classifying mild words.

Added

2026-09-24