Bag-of-Words vs. Graph vs. Sequence in Text Classification: Questioning the Necessity of Text-Graphs and the Surprising Strength of a Wide MLP
Lukas GalkeAnsgar Scherp
Demonstrates that a simple, wide bag-of-words multi-layer perceptron outperforms complex graph neural networks like TextGCN in inductive text classification while offering significantly faster training and inference than Transformer models on long sequences.
Text categorization—assigning topical labels to documents, articles, or social media posts—is a foundational capability in modern natural language processing. In recent years, academic research has increasingly favored complex graph-based architectures, such as graph convolutional networks that construct large synthetic word-and-document graphs across entire corpora. However, these intricate structures present substantial deployment friction, including massive memory footprints and difficulty classifying new, unseen documents without retraining. Meanwhile, straightforward multi-layer perceptrons (MLPs)—basic feedforward neural networks—have largely been overlooked as competitive baselines.
The article evaluates whether the architectural complexity of synthetic text graphs is genuinely necessary for competitive text categorization. It systematically benchmarks simple bag-of-words feedforward models against modern graph neural networks and large pretrained sequence-based Transformer models across standard inductive and transductive classification environments.
To conduct this evaluation, the authors assessed 16 methods across three core architectural families: word-count models, graph models, and sequence models. They ran direct experiments using a custom wide MLP (featuring a single wide hidden layer of 1,024 units), full BERT, and lightweight DistilBERT, comparing their performance against published benchmarks across five standard datasets spanning newsgroups, medical abstracts, and movie reviews (20ng, R8, R52, Ohsumed, and MR).
The investigation produced several critical findings. First, in realistic inductive settings where incoming test documents are previously unseen, the simple wide MLP outperforms prominent graph-based models like TextGCN and HeteGCN, while remaining competitive with more complex hypergraph methods. Second, pretrained sequence Transformers set a new overall performance benchmark: BERT attained the highest overall accuracy, outperforming the leading graph model by 8 points on sentiment analysis and by 0.5 to 1.5 points on topical datasets, closely tracked by DistilBERT. Third, the wide MLP demonstrated superior operational efficiency, requiring roughly 31.3 million parameters—about half that of DistilBERT and a fraction of BERT—while achieving training runtimes an order of magnitude faster per epoch on long-text datasets.
These findings indicate that creating synthetic graphs from text documents offers little to no practical advantage over simpler architectures for inductive text classification. Pretrained word embeddings such as GloVe also proved less effective in MLPs than training directly on word-frequency representations, as wide hidden layers naturally avoid embedding-dimension bottlenecks. For real-world deployments, organizations can achieve state-of-the-art accuracy using distilled Transformers, or choose wide MLPs to obtain near-equivalent performance with markedly lower computational overhead, reduced operational latency, and zero graph-construction costs.
For practical implementation, organizations facing strict computational or latency constraints should deploy a single-layer wide MLP using modern subword tokenizers and standard dropout regularization. When absolute accuracy is paramount and computing resources permit, fine-tuning lightweight sequence models like DistilBERT represents the optimal choice. Future evaluations should establish the wide MLP as an essential baseline before adopting more intricate deep learning architectures.
Confidence in these comparative findings is high due to the standardized datasets, rigorous train-test splits, and consistent multi-run evaluations. However, practitioners should note that the scope was restricted to single-label English text categorization tasks. Caution is advised when generalizing these conclusions directly to languages requiring specialized structural modeling or to complex classification domains such as multi-label, hierarchical, or few-shot categorization.
- Paper: Graph Convolutional Networks for Text Classification, Liang Yao et al. (2018). Read TextGCN first to understand the graph-based text-classification benchmark that the source directly challenges.
- Paper: Baselines and Bigrams: Simple, Good Sentiment and Topic Classification, Sida I. Wang et al. (2012). Its carefully tested Naive Bayes and SVM baselines provide the simple-classifier context for the source’s case for strong, lightweight models.
- Paper: BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, Jacob Devlin et al. (2019). BERT establishes the pretrained Transformer classification approach that the source evaluates against bag-of-words and graph models.
- Paper: Semi-Supervised Classification with Graph Convolutional Networks, Thomas N. Kipf et al. (2017). This foundational GCN paper explains the graph-convolution method underlying graph architectures later adapted to text classification.
No sufficiently relevant recommendations were found.
