Bag-of-Words vs. Graph vs. Sequence in Text Classification: Questioning the Necessity of Text-Graphs and the Surprising Strength of a Wide MLP

Lukas GalkeAnsgar Scherp

article2022ACL54 citations

Demonstrates that a simple, wide bag-of-words multi-layer perceptron outperforms complex graph neural networks like TextGCN in inductive text classification while offering significantly faster training and inference than Transformer models on long sequences.

Listen

Text categorization—assigning topical labels to documents, articles, or social media posts—is a foundational capability in modern natural language processing. In recent years, academic research has increasingly favored complex graph-based architectures, such as graph convolutional networks that construct large synthetic word-and-document graphs across entire corpora. However, these intricate structures present substantial deployment friction, including massive memory footprints and difficulty classifying new, unseen documents without retraining. Meanwhile, straightforward multi-layer perceptrons (MLPs)—basic feedforward neural networks—have largely been overlooked as competitive baselines.

The article evaluates whether the architectural complexity of synthetic text graphs is genuinely necessary for competitive text categorization. It systematically benchmarks simple bag-of-words feedforward models against modern graph neural networks and large pretrained sequence-based Transformer models across standard inductive and transductive classification environments.

To conduct this evaluation, the authors assessed 16 methods across three core architectural families: word-count models, graph models, and sequence models. They ran direct experiments using a custom wide MLP (featuring a single wide hidden layer of 1,024 units), full BERT, and lightweight DistilBERT, comparing their performance against published benchmarks across five standard datasets spanning newsgroups, medical abstracts, and movie reviews (20ng, R8, R52, Ohsumed, and MR).

The investigation produced several critical findings. First, in realistic inductive settings where incoming test documents are previously unseen, the simple wide MLP outperforms prominent graph-based models like TextGCN and HeteGCN, while remaining competitive with more complex hypergraph methods. Second, pretrained sequence Transformers set a new overall performance benchmark: BERT attained the highest overall accuracy, outperforming the leading graph model by 8 points on sentiment analysis and by 0.5 to 1.5 points on topical datasets, closely tracked by DistilBERT. Third, the wide MLP demonstrated superior operational efficiency, requiring roughly 31.3 million parameters—about half that of DistilBERT and a fraction of BERT—while achieving training runtimes an order of magnitude faster per epoch on long-text datasets.

These findings indicate that creating synthetic graphs from text documents offers little to no practical advantage over simpler architectures for inductive text classification. Pretrained word embeddings such as GloVe also proved less effective in MLPs than training directly on word-frequency representations, as wide hidden layers naturally avoid embedding-dimension bottlenecks. For real-world deployments, organizations can achieve state-of-the-art accuracy using distilled Transformers, or choose wide MLPs to obtain near-equivalent performance with markedly lower computational overhead, reduced operational latency, and zero graph-construction costs.

For practical implementation, organizations facing strict computational or latency constraints should deploy a single-layer wide MLP using modern subword tokenizers and standard dropout regularization. When absolute accuracy is paramount and computing resources permit, fine-tuning lightweight sequence models like DistilBERT represents the optimal choice. Future evaluations should establish the wide MLP as an essential baseline before adopting more intricate deep learning architectures.

Confidence in these comparative findings is high due to the standardized datasets, rigorous train-test splits, and consistent multi-run evaluations. However, practitioners should note that the scope was restricted to single-label English text categorization tasks. Caution is advised when generalizing these conclusions directly to languages requiring specialized structural modeling or to complex classification domains such as multi-label, hierarchical, or few-shot categorization.

arXiv: 2109.03777

No sufficiently relevant recommendations were found.

Cover for Bag-of-Words vs. Graph vs. Sequence in Text Classification: Questioning the Necessity of Text-Graphs and the Surprising Strength of a Wide MLP

Abstract

Graph neural networks have triggered a resurgence of graph-based text classification methods, defining today’s state of the art. We show that a wide multi-layer perceptron (MLP) using a Bag-of-Words (BoW) outperforms the recent graph-based models TextGCN and HeteGCN in an inductive text classification setting and is comparable with HyperGAT. Moreover, we fine-tune a sequence-based BERT and a lightweight DistilBERT model, which both outperform all state-of-the-art models. These results question the importance of synthetic graphs used in modern text classifiers. In terms of efficiency, DistilBERT is still twice as large as our BoW-based wide MLP, while graph-based models like TextGCN require setting up an O(N²) graph, where N is the vocabulary plus corpus size. Finally, since Transformers need to compute O(L²) attention weights with sequence length L, the MLP models show higher training and inference speeds on datasets with long sequences.

Table of Contents

  • 1 Introduction
  • 2 Literature on Text Categorization
  • 3 Models for Text Categorization
  • 3.1 BoW-Based Text Categorization
  • 3.2 Graph-Based Text Categorization
  • 3.3 Sequence-Based Text Categorization
  • 4 Experimental Apparatus
  • 4.1 Datasets
  • 4.2 Inductive and Transductive Setups
  • 4.3 Procedure and Hyperparameter Settings
  • 4.4 Measures
  • 5 Results
  • 5.1 Effectiveness
  • 5.2 Efficiency
  • 6 Discussion
  • 7 Conclusion
  • Ethical Considerations
  • References
  • A Practical Guidelines for Designing a WideMLP
  • B Connection between BoW-MLP and TextGCN
  • C Equivalence of Micro-F1 and Accuracy in Multiclass Classification

Knowls

  1. Knowl 1 — Inductive and transductive accuracy comparison

    data/table

    The comparison evaluates single-label text classification accuracy (percent) on the same five benchmark datasets. “Inductive” results use models that do not have access to test documents during preprocessing; “transductive” graph results use the test documents in the graph construction or preprocessing. Parenthesized values are standard deviations where reported; dashes indicate unavailable results. WideMLP is competitive with or better than several inductive graph models, while BERT and DistilBERT score strongly across the datasets. HyperGAT is the strongest graph comparator in most of the paper’s discussion; DADGNN has higher reported scores than BERT on R8 and higher than the other listed models on R52 among the graph results. The table also shows that adding depth or pretrained GloVe embeddings does not consistently improve the WideMLP results.

    Setting Model 20ng R8 R52 ohsumed MR Provenance
    Inductive Logistic regression 83.70 93.33 90.65 61.14 76.28 Ragesh et al. (2021)
    Inductive SWEM 85.16 (0.29) 95.32 (0.26) 92.94 (0.24) 63.12 (0.55) 76.65 (0.63) Ding et al. (2020)
    Inductive fastText 79.38 (0.30) 96.13 (0.21) 92.81 (0.09) 57.70 (0.49) 75.14 (0.20) Ding et al. (2020)
    Inductive TF-IDF + WideMLP 84.20 (0.16) 97.08 (0.16) 93.67 (0.23) 66.06 (0.29) 76.32 (0.17) Own experiment
    Inductive WideMLP 83.31 (0.22) 97.27 (0.12) 93.89 (0.16) 63.95 (0.13) 76.72 (0.26) Own experiment
    Inductive WideMLP-2 81.02 (0.23) 96.61 (1.22) 93.98 (0.23) 61.71 (0.33) 75.91 (0.51) Own experiment
    Inductive GloVe + WideMLP 76.80 (0.11) 96.44 (0.08) 93.58 (0.06) 61.36 (0.22) 75.96 (0.17) Own experiment
    Inductive GloVe + WideMLP-2 76.33 (0.18) 96.50 (0.14) 93.19 (0.11) 61.65 (0.27) 75.72 (0.45) Own experiment
    Inductive TextGCN 80.88 (0.54) 94.00 (0.40) 89.39 (0.38) 56.32 (1.36) 74.60 (0.43) Ragesh et al. (2021)
    Inductive HeteGCN 84.59 (0.14) 97.17 (0.33) 93.89 (0.45) 63.79 (0.80) 75.62 (0.26) Ragesh et al. (2021)
    Inductive HyperGAT 86.62 (0.16) 97.07 (0.23) 94.98 (0.27) 69.90 (0.34) 78.32 (0.27) Ragesh et al. (2021)
    Inductive DADGNN – 98.15 (0.16) 95.16 (0.22) – 78.64 (0.29) Liu et al. (2021b)
    Inductive Pretrained LSTM 75.43 (1.72) 96.09 (0.19) 90.48 (0.86) 51.10 (1.50) 77.33 (0.89) Ding et al. (2020)
    Inductive DistilBERT 86.24 (0.26) 97.89 (0.15) 95.34 (0.08) 69.08 (0.60) 85.10 (0.33) Own experiment
    Inductive BERT 87.21 (0.18) 98.03 (0.24) 96.17 (0.33) 71.46 (0.54) 86.61 (0.38) Own experiment
    Inductive BERT without position embeddings 81.47 (0.49) 97.39 (0.20) 94.70 (0.27) 65.18 (1.53) 80.35 (0.20) Own experiment
    Inductive BERT with shuffled augmentation 86.46 (0.42) 98.07 (0.21) 96.48 (0.18) 70.94 (0.60) 86.23 (0.33) Own experiment
    Transductive TextGCN 86.34 97.07 93.56 68.36 76.74 Yao et al. (2019)
    Transductive SGC 88.5 (0.1) 97.2 (0.1) 94.0 (0.2) 68.5 (0.3) 75.9 (0.3) Wu et al. (2019)
    Transductive TensorGCN 87.74 98.04 95.05 70.11 77.91 Liu et al. (2020)
    Transductive HeteGCN 87.15 (0.15) 97.24 (0.51) 94.35 (0.25) 68.11 (0.70) 76.71 (0.33) Ragesh et al. (2021)
  2. Knowl 2 — WideMLP design and training recipe

    model/method

    The paper’s WideMLP classifies a document from a bag-of-words representation, without using token order or a corpus-induced graph. The main model has one hidden layer of 1,024 ReLU units followed by an output layer; WideMLP-2 adds a second 1,024-unit ReLU hidden layer. Inputs tested were unweighted bag-of-words counts, TF-IDF-weighted counts, and averaged GloVe representations. For frozen GloVe inputs, dropout and ReLU are not applied to the frozen embeddings, but are applied to subsequent layers. The authors use BERT’s uncased WordPiece tokenizer for the text representations.

    Training minimizes cross-entropy with Adam at its default learning rate of 10−310^{-3}, a linearly decaying learning-rate schedule, 100 epochs, batch size 16, and dropout ratio 0.5. Dropout is applied after hidden layers, including after the initial embedding layer. Each experiment is repeated with five random parameter initializations. The authors report that the MLP was not very sensitive to hyperparameter choices in their preliminary experiments. They also note that the wide input layer can be implemented as an embedding lookup followed by aggregation, avoiding a large dense matrix multiplication.

  3. Knowl 3 — Efficiency: parameter counts and measured runtimes

    data/table

    The parameter counts compare the WideMLP with the Transformer baselines; the runtime measurements include training and inference, average five runs, and are rounded to minutes. Runtime experiments used one NVIDIA A100-SXM4-40GB GPU. WideMLP has 31.3M parameters, compared with 66M for DistilBERT and 110M for BERT. The GloVe parameter totals are large, but most of those parameters are frozen. In the measured total runtimes, WideMLP variants are faster than the Transformers on the longer datasets; Transformers are faster on MR, whose texts are much shorter. The authors also report that DistilBERT is about twice as fast as BERT and that WideMLP variants have an order-of-magnitude advantage over DistilBERT in average runtime per epoch. Their complexity discussion notes that standard Transformer attention requires O(L2)O(L^2) weights for sequence length LL, while a word-document graph requires O(N2)O(N^2) space for NN equal to the vocabulary size plus the corpus size.

    Model Parameters
    WideMLP 31.3M
    WideMLP-2 32.3M
    GloVe + WideMLP 575.2M frozen + 0.3M
    GloVe + WideMLP-2 575.2M frozen + 1.3M
    DistilBERT 66M
    BERT 110M
    Model Epochs 20ng R8 R52 ohsumed MR
    WideMLP 100 7 min 3 min 4 min 3 min 4 min
    TF-IDF + WideMLP 100 9 min 4 min 4 min 3 min 4 min
    WideMLP-2 100 9 min 5 min 5 min 3 min 6 min
    GloVe + WideMLP 100 6 min 3 min 4 min 3 min 4 min
    GloVe + WideMLP-2 100 6 min 4 min 4 min 3 min 4 min
    DistilBERT 10 8 min 4 min 5 min 3 min 1 min
    BERT 10 15 min 7 min 8 min 5 min 2 min
  4. Knowl 4 — Benchmark datasets and evaluation protocol

    experimental setup

    The experiments use the TextGCN train-test splits for 20ng, R8, R52, ohsumed, and MR. 20ng is the by-date version of the 20 Newsgroups corpus; R8 and R52 are Reuters 21578 subsets; ohsumed contains medical abstracts; MR is a binary movie-review sentiment dataset using the split of Tang et al. The listed sequence lengths are mean word counts and standard deviations. Accuracy is the evaluation metric, equivalent to micro-F1 for single-label classification. The authors repeat their own experiments five times and report the mean and standard deviation.

    Dataset NN Train Test Classes Mean length (SD) Task/domain
    20ng 18,846 11,314 7,532 20 551 (2,047) Newsgroup categorization
    R8 7,674 5,485 2,189 8 119 (128) Reuters categorization
    R52 9,100 6,532 2,568 52 126 (133) Reuters categorization
    oh sumed 7,400 3,357 4,043 23 285 (123) Medical abstracts
    MR 10,662 7,108 3,554 2 25 (11) Movie-review sentiment

    In an inductive evaluation, test documents are unavailable during preprocessing and are encountered only at test time. In a transductive evaluation, test documents are visible and may be used during preprocessing, as in the original TextGCN setup. Bag-of-words and sequence models are inductive by construction; the paper reports graph-model results in both settings where the literature provides them. BERT inputs are truncated to 512 tokens.

  5. Knowl 5 — BERT fine-tuning and word-order ablations

    experimental setup

    BERT and DistilBERT are fine-tuned for 10 epochs with a linearly decaying learning rate starting at 5×10−55\times10^{-5}. The effective batch size is 128, obtained by gradient accumulation over eight batches of 16, and sequences are truncated to 512 tokens. To test the role of position information, the authors set BERT’s position embeddings to zero and disable their gradients, making the input representation position-agnostic. A second ablation augments training by shuffling each sequence, thereby doubling the number of training examples. In the reported results, removing position embeddings lowers BERT accuracy on all five datasets, although performance does not collapse; shuffled-sequence augmentation has no consistent direction of effect. The exact accuracies for these conditions are included in the comparison table.

  6. Knowl 6 — TextGCN’s first layer as embedding aggregation

    model/method

    For TextGCN with identity node features, the first graph-convolution layer can be viewed as adjacency-weighted aggregation of learned node embeddings. Let ii be a graph node, N(i)\mathcal{N}(i) its neighbors, aija_{ij} the graph’s edge weight from node jj to node ii (with self-loops inserted), and EjE_j the learned embedding associated with node jj. The first-layer representation is

    hi=∑j∈N(i)aijEj.h_i=\sum_{j\in\mathcal{N}(i)}a_{ij}E_j.

    Because nodes include both words and documents, this aggregation combines word embeddings with document embeddings. The paper characterizes a single-layer TextGCN as equivalent in this respect to an embedding-based bag-of-words MLP, apart from the document embedding and graph-derived weights. A second TextGCN layer can additionally incorporate embeddings of other documents connected through shared words. Thus, the graph model’s corpus-level structure and document-node embeddings distinguish it from an inductive MLP that sees only a new document’s bag of words.

  7. Knowl 7 — Properties of the compared model families

    definition

    The paper organizes text classifiers by whether they construct a synthetic graph, represent word position, accept arbitrary-length text, and support inductive prediction. Bag-of-words models use no graph and discard position; they accept arbitrary-length inputs and are inductive. TextGCN uses a synthetic graph, discards word order, accepts arbitrary-length text, and is not inductive in its original transductive form. TensorGCN also uses a graph, represents position, accepts arbitrary lengths, and is not inductive in the comparison. HeteGCN and HyperGAT use synthetic graphs but are treated as position-unaware, arbitrary-length, inductive models. RNN/CNN sequence models use position without synthetic graphs and support arbitrary lengths; BERT/DistilBERT use position without synthetic graphs and are inductive, but have a fixed maximum input length. This distinction matters because original transductive graph methods require test-corpus information before prediction, whereas the paper’s central comparison is framed around inductive classification.

  8. Knowl 8 — Scope and limits of the evidence

    limitation

    The empirical evaluation covers five English-language datasets and single-label classification, spanning topical categorization and movie-review sentiment. The paper therefore does not establish that the same ranking holds for other languages, multi-label or hierarchical classification, or few-shot settings; these are identified as directions for further evaluation. Some graph-model scores are taken from prior publications rather than rerun by the authors, although the study says it checked train-test splits, class counts, preprocessing, training procedures, and reported metrics for comparability. Runtime comparisons also depend on training parameters such as batch size, so the reported timings are specific to the stated setup rather than hardware-independent efficiency guarantees.

Coverage note — The literature-survey details and the appendix proof that accuracy equals micro-F1 for single-label classification are omitted because they do not add a distinct result beyond the benchmark setup and contributed model comparisons.

References

  1. 1.Markus Bayer, Marc-André Kaufhold, and Christian Reuter. 2021. A survey on data augmentation for text classification. ArXiv preprint, abs/2107.03158.
  2. 2.Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’21, pages 610–623, New York, NY, USA. Association for Computing Machinery.
  3. 3.David M. Blei, Andrew Y. Ng, and Michael I. Jordan. 2001. Latent dirichlet allocation. In Advances in Neural Information Processing Systems 14 [Neural Information Processing Systems: Natural and Synthetic, NIPS 2001, December 3-8, 2001, Vancouver, British Columbia, Canada], pages 601–608. MIT Press.
  4. 4.Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017. Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5:135–146.
  5. 5.Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, et al. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual.
  6. 6.Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018. What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2126–2136, Melbourne, Australia. Association for Computational Linguistics.
  7. 7.George Cybenko. 1989. Approximation by superpositions of a sigmoidal function. Math. Control. Signals Syst., 2(4):303–314.
  8. 8.Maurizio Ferrari Dacrema, Paolo Cremonesi, and Dietmar Jannach. 2019. Are we really making much progress? A worrying analysis of recent neural recommendation approaches. In Proceedings of the 13th ACM Conference on Recommender Systems, RecSys 2019, Copenhagen, Denmark, September 16-20, 2019, pages 101–109. ACM.
  9. 9.Mostafa Dehghani, Anurag Arnab, Lucas Beyer, Ashish Vaswani, and Yi Tay. 2021. The efficiency misnomer. ArXiv preprint, abs/2110.12894.
  10. 10.Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  11. 11.Kaize Ding, Jianling Wang, Jundong Li, Dingcheng Li, and Huan Liu. 2020. Be more with less: Hypergraph attention networks for inductive text classification. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4927–4936, Online. Association for Computational Linguistics.
  12. 12.Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net.
  13. 13.Matthias Fey, Jan Eric Lenssen, Frank Weichert, and Jure Leskovec. 2021. GNNAutoScale: Scalable and expressive graph neural networks via historical embeddings. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 3294–3304. PMLR.
  14. 14.Quentin Fournier, Gaétan Marceau Caron, and Daniel Aloise. 2021. A practical survey on faster and lighter transformers. ArXiv preprint, abs/2103.14636.
  15. 15.Lukas Galke, Florian Mai, Alan Schelten, Dennis Brunsch, and Ansgar Scherp. 2017. Using titles vs. full-text as source for automated semantic document annotation. In Proceedings of the Knowledge Capture Conference, K-CAP 2017, Austin, TX, USA, December 4-6, 2017, pages 20:1–20:4. ACM.
  16. 16.Santiago González-Carvajal and Eduardo C. Garrido-Merchán. 2020. Comparing BERT against traditional machine learning text classification. ArXiv preprint, abs/2005.13012.
  17. 17.William L. Hamilton. 2020. Graph Representation Learning. Synthesis Lectures on Artificial Intelligence and Machine Learning. Morgan & Claypool Publishers.
  18. 18.Mohit Iyyer, Varun Manjunatha, Jordan Boyd-Graber, and Hal Daumé III. 2015. Deep unordered composition rivals syntactic methods for text classification. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1681–1691, Beijing, China. Association for Computational Linguistics.
  19. 19.Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2020. TinyBERT: Distilling BERT for natural language understanding. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 4163–4174, Online. Association for Computational Linguistics.
  20. 20.Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2017. Bag of tricks for efficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 427–431, Valencia, Spain. Association for Computational Linguistics.
  21. 21.Ammar Ismael Kadhim. 2019. Survey on supervised machine learning techniques for automatic text classification. Artif. Intell. Rev., 52(1):273–292.
  22. 22.Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom. 2014. A convolutional neural network for modelling sentences. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 655–665, Baltimore, Maryland. Association for Computational Linguistics.
  23. 23.Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings.
  24. 24.Thomas N. Kipf and Max Welling. 2017. Semi-supervised classification with graph convolutional networks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings. OpenReview.net.
  25. 25.Kamran Kowsari, Kiana Jafari Meimandi, Mojtaba Heidarysafa, Sanjana Mendu, Laura E. Barnes, and Donald E. Brown. 2019. Text classification algorithms: A survey. Inf., 10(4):150.
  26. 26.Siwei Lai, Liheng Xu, Kang Liu, and Jun Zhao. 2015. Recurrent convolutional neural networks for text classification. In Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, January 25-30, 2015, Austin, Texas, USA, pages 2267–2273. AAAI Press.
  27. 27.Qian Li, Hao Peng, Jianxin Li, Congying Xia, Renyu Yang, Lichao Sun, Philip S. Yu, and Lifang He. 2020. A survey on text classification: From shallow to deep learning. ArXiv preprint, abs/2008.00364.
  28. 28.Hanxiao Liu, Zihang Dai, David R. So, and Quoc V. Le. 2021a. Pay attention to MLPs. ArXiv preprint, abs/2105.08050.
  29. 29.Xien Liu, Xinxin You, Xiao Zhang, Ji Wu, and Ping Lv. 2020. Tensor graph convolutional networks for text classification. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, IAAI 2020, The Tenth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2020, New York, NY, USA, February 7-12, 2020, pages 8409–8416. AAAI Press.
  30. 30.Yonghao Liu, Renchu Guan, Fausto Giunchiglia, Yanchun Liang, and Xiaoyue Feng. 2021b. Deep attention diffusion graph neural networks for text classification. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8142–8152, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  31. 31.Haonan Lu, Seth H. Huang, Tian Ye, and Xiuyan Guo. 2019. Graph star net for generalized multi-task learning. ArXiv preprint, abs/1906.12330.
  32. 32.Shengfei Lyu and Jiaqi Liu. 2020. Combine convolution with recurrent networks for text classification. ArXiv preprint, abs/2006.15795.
  33. 33.Florian Mai, Lukas Galke, and Ansgar Scherp. 2018. Using deep learning for title-based semantic subject indexing to reach competitive performance to full-text. In Proceedings of the 18th ACM/IEEE on Joint Conference on Digital Libraries, JCDL 2018, Fort Worth, TX, USA, June 03-07, 2018, pages 169–178. ACM.
  34. 34.Luke Melas-Kyriazi. 2021. Do you even need attention? A stack of feed-forward layers does surprisingly well on ImageNet. ArXiv preprint, abs/2105.02723.
  35. 35.Tomás Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems 26: 27th Annual Conference on Neural Information Processing Systems 2013. Proceedings of a meeting held December 5-8, 2013, Lake Tahoe, Nevada, United States, pages 3111–3119.
  36. 36.Shervin Minaee, Nal Kalchbrenner, Erik Cambria, Narjes Nikzad, Meysam Chenaghlu, and Jianfeng Gao. 2021. Deep learning-based text classification: A comprehensive review. ACM Comput. Surv., 54(3):62:1–62:40.
  37. 37.Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever. 2020. Deep double descent: Where bigger models and more data hurt. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net.
  38. 38.Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro. 2018. Towards understanding the role of over-parametrization in generalization of neural networks. ArXiv preprint, abs/1805.12076.
  39. 39.Yukio Ohsawa, Nels E. Benson, and Masahiko Yachida. 1998. Keygraph: Automatic indexing by co-occurrence graph based on building construction metaphor. In Proceedings of the IEEE Forum on Research and Technology Advances in Digital Libraries, IEEE ADL ’98, Santa Barbara, California, USA, April 22-24, 1998, pages 12–18. IEEE Computer Society.
  40. 40.Bo Pang and Lillian Lee. 2005. Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales. In Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL’05), pages 115–124, Ann Arbor, Michigan. Association for Computational Linguistics.
  41. 41.Hao Peng, Jianxin Li, Yu He, Yaopeng Liu, Mengjiao Bao, Lihong Wang, Yangqiu Song, and Qiang Yang. 2018. Large-scale hierarchical text classification with recursively regularized deep graph-cnn. In Proceedings of the 2018 World Wide Web Conference on World Wide Web, WWW 2018, Lyon, France, April 23-27, 2018, pages 1063–1072. ACM.
  42. 42.Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014. GloVe: Global vectors for word representation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 1532–1543, Doha, Qatar. Association for Computational Linguistics.
  43. 43.Rahul Ragesh, Sundararajan Sellamanickam, Arun Iyer, Ramakrishna Bairi, and Vijay Lingam. 2021. HeteGCN: Heterogeneous graph convolutional networks for text classification. In WSDM ’21, The Fourteenth ACM International Conference on Web Search and Data Mining, Virtual Event, Israel, March 8-12, 2021, pages 860–868. ACM.
  44. 44.Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas. Association for Computational Linguistics.
  45. 45.Stuart Rose, Dave Engel, Nick Cramer, and Wendy Cowley. 2010. Automatic keyword extraction from individual documents. In Text Mining, chapter 1, pages 1–20. John Wiley & Sons, Ltd.
  46. 46.Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter. ArXiv preprint, abs/1910.01108.
  47. 47.Oleksandr Shchur, Maximilian Mumme, Aleksandar Bojchevski, and Stephan Günnemann. 2018. Pitfalls of graph neural network evaluation. ArXiv preprint, abs/1811.05868.
  48. 48.Dinghan Shen, Guoyin Wang, Wenlin Wang, Martin Renqiang Min, Qinliang Su, Yizhe Zhang, Chunyan Li, Ricardo Henao, and Lawrence Carin. 2018. Baseline needs more love: On simple word-embedding-based models and associated pooling mechanisms. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 440–450, Melbourne, Australia. Association for Computational Linguistics.
  49. 49.Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela. 2021. Masked language modeling and the distributional hypothesis: Order word matters pre-training for little. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2888–2913, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  50. 50.Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou. 2020. MobileBERT: a compact task-agnostic BERT for resource-limited devices. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 2158–2170, Online. Association for Computational Linguistics.
  51. 51.Jian Tang, Meng Qu, and Qiaozhu Mei. 2015. PTE: predictive text embedding through large-scale heterogeneous text networks. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Sydney, NSW, Australia, August 10-13, 2015, pages 1165–1174. ACM.
  52. 52.Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2020. Efficient transformers: A survey. ArXiv preprint, abs/2009.06732.
  53. 53.Ilya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy. 2021. MLP-Mixer: An all-MLP architecture for vision. ArXiv preprint, abs/2105.01601.
  54. 54.Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008.
  55. 55.Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. 2018. Graph attention networks. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net.
  56. 56.Ishwar Venugopal, Jessica Töllich, Michael Fairbank, and Ansgar Scherp. 2021. A comparison of deep-learning methods for analysing and predicting business processes. In International Joint Conference on Neural Networks, IJCNN 2021, Shenzhen, China, July 18-22, 2021, pages 1–8. IEEE.
  57. 57.Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019a. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net.
  58. 58.Ruishuang Wang, Zhao Li, Jian Cao, Tong Chen, and Lei Wang. 2019b. Convolutional recurrent neural networks for text classification. In International Joint Conference on Neural Networks, IJCNN 2019 Budapest, Hungary, July 14-19, 2019, pages 1–6. IEEE.
  59. 59.Yequan Wang, Aixin Sun, Jialong Han, Ying Liu, and Xiaoyan Zhu. 2018. Sentiment analysis by capsules. In Proceedings of the 2018 World Wide Web Conference on World Wide Web, WWW 2018, Lyon, France, April 23-27, 2018, pages 1165–1174. ACM.
  60. 60.Felix Wu, Amauri H. Souza Jr., Tianyi Zhang, Christopher Fifty, Tao Yu, and Kilian Q. Weinberger. 2019. Simplifying graph convolutional networks. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA, volume 97 of Proceedings of Machine Learning Research, pages 6861–6871. PMLR.
  61. 61.Lingfei Wu, Yu Chen, Kai Shen, Xiaojie Guo, Hanning Gao, Shucheng Li, Jian Pei, and Bo Long. 2021. Graph neural networks for natural language processing: A survey. ArXiv preprint, abs/2106.06090.
  62. 62.Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, et al. 2016. Google’s neural machine translation system: Bridging the gap between human and machine translation. ArXiv preprint, abs/1609.08144.
  63. 63.Huiru Xiao, Xin Liu, and Yangqiu Song. 2019. Efficient path prediction for semi-supervised and weakly supervised hierarchical text classification. In The World Wide Web Conference, WWW 2019, San Francisco, CA, USA, May 13-17, 2019, pages 3370–3376. ACM.
  64. 64.Liang Yao, Chengsheng Mao, and Yuan Luo. 2019. Graph convolutional networks for text classification. In The Thirty-Third AAAI Conference on Artificial Intelligence, AAAI 2019, The Thirty-First Innovative Applications of Artificial Intelligence Conference, IAAI 2019, The Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, EAAI 2019, Honolulu, Hawaii, USA, January 27 - February 1, 2019, pages 7370–7377. AAAI Press.
  65. 65.Rahul Yedida, Xueqi Yang, and Tim Menzies. 2021. When SIMPLE is better than complex: A case study on deep learning for predicting Bugzilla issue close time. ArXiv preprint, abs/2101.06319.
  66. 66.Dell Zhang, Jun Wang, Emine Yilmaz, Xiaoling Wang, and Yuxin Zhou. 2016. Bayesian performance comparison of text classifiers. In Proceedings of the 39th International ACM SIGIR conference on Research and Development in Information Retrieval, SIGIR 2016, Pisa, Italy, July 17-21, 2016, pages 15–24. ACM.
  67. 67.Lu Zhang, Jiandong Ding, Yi Xu, Yingyao Liu, and Shuigeng Zhou. 2021. Weakly-supervised text classification based on keyword graph. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 2803–2813, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  68. 68.Peng Zhou, Zhenyu Qi, Suncong Zheng, Jiaming Xu, Hongyun Bao, and Bo Xu. 2016. Text classification improved by integrating bidirectional LSTM with two-dimensional max pooling. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, pages 3485–3495, Osaka, Japan. The COLING 2016 Organizing Committee.
  69. 69.Xujuan Zhou, Raj Gururajan, Yuefeng Li, Revathi Venkataraman, Xiaohui Tao, Ghazal Bargshady, Prabal Datta Barua, and Srinivas Kondalsamy-Chennakesavan. 2020. A survey on text classification and its applications. Web Intell., 18(3):205–216.

Citation

MLA
Galke, L., and A. Scherp. “Bag-of-Words Vs. Graph Vs. Sequence in Text Classification: Questioning the Necessity of Text-Graphs and the Surprising Strength of a Wide MLP”. arXiv, 2021, http://arxiv.org/abs/2109.03777v3.
APA
Galke, L., & Scherp, A. (2021). Bag-of-Words vs. Graph vs. Sequence in Text Classification: Questioning the Necessity of Text-Graphs and the Surprising Strength of a Wide MLP. arXiv. http://arxiv.org/abs/2109.03777v3
Chicago
Galke, L., and A. Scherp. 2021. “Bag-of-Words Vs. Graph Vs. Sequence in Text Classification: Questioning the Necessity of Text-Graphs and the Surprising Strength of a Wide MLP”. arXiv. http://arxiv.org/abs/2109.03777v3.
Harvard
Galke, L. and Scherp, A. (2021) “Bag-of-Words vs. Graph vs. Sequence in Text Classification: Questioning the Necessity of Text-Graphs and the Surprising Strength of a Wide MLP”, arXiv [Preprint]. Available at: http://arxiv.org/abs/2109.03777v3.
Vancouver
1. Galke L, Scherp A (2021) Bag-of-Words vs. Graph vs. Sequence in Text Classification: Questioning the Necessity of Text-Graphs and the Surprising Strength of a Wide MLP. arXiv

BibTeX

@article{galke2021bag,
  title = {Bag-of-Words vs. Graph vs. Sequence in Text Classification: Questioning the Necessity of Text-Graphs and the Surprising Strength of a Wide MLP},
  author = {Galke, Lukas and Scherp, Ansgar},
  year = {2021},
  journal = {arXiv},
  url = {http://arxiv.org/abs/2109.03777v3},
  eprint = {2109.03777}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/