“Liar, Liar Pants on Fire”: A New Benchmark Dataset for Fake News Detection

William Yang Wang

article2017ACL1,740 citations

Introduces the LIAR benchmark dataset of 12,800 fact-checked statements from PolitiFact alongside a hybrid neural network architecture that integrates contextual metadata with text to improve automated fake news detection.

Listen

The rapid spread of misinformation poses severe risks to political stability, public discourse, and safety. Addressing this problem with automated statistical tools has been severely limited by a lack of benchmark data, as previous fact-checking datasets contain only a few hundred examples or rely on artificial, crowdsourced text that fails to reflect real-world political speech.

The article demonstrates the creation of a large-scale, benchmark dataset for automated fake news detection and evaluates how well text-based and hybrid machine learning models classify statements into fine-grained categories of truthfulness.

The author compiled the LIAR dataset, which contains 12,836 real-world short statements sourced from the fact-checking website PolitiFact between 2007 and 2016. This collection is an order of magnitude larger than previous resources. Each statement includes context, speaker metadata (such as political party, job, home state, and historical accuracy record), a detailed analysis report, and one of six truthfulness ratings. The evaluation compared standard text classification algorithms against a hybrid neural network architecture designed to integrate textual claims with speaker metadata.

The key findings reveal that automatic detection of nuanced misinformation remains difficult. A baseline predicting the most common class achieved 20.8% test accuracy across the six-category task. Standard machine learning models improved accuracy to roughly 25%. A deep learning model analyzing text alone reached 27.0% accuracy, significantly outperforming standard baselines. Ultimately, the best performance was achieved by the hybrid model that combined the text statement with all available metadata, reaching an accuracy of 27.4% on the test set.

These results indicate that while metadata provides a measurable boost, text analysis and metadata alone are insufficient for reliable, high-stakes fake news detection. Because fine-grained truthfulness classification yields low absolute accuracy, organizations cannot deploy automated filters in isolation without risking high false-positive or false-negative rates. The presence of rich supporting evidence in the dataset suggests that automated systems must evolve beyond surface-level language to incorporate external knowledge.

Decision-makers and researchers should use this benchmark dataset to develop more advanced fact-checking tools that connect automated models directly to external knowledge bases and reference documents. Future efforts should explore related areas such as argument mining, rumor detection, and stance classification to build robust verification systems.

Confidence in the benchmark's quality is reinforced by an inter-annotator agreement rate of 0.82 between the author and original professional fact-checkers. However, leaders should note that the dataset reflects political statements specific to United States politics, meaning performance and findings may vary when applied to other domains, regions, or languages.

arXiv: 1705.00648
  • Paper: Information credibility on twitter, Carlos Castillo et al. (2011). This foundational study demonstrates how supervised classification and metadata features can be leveraged to automatically determine information credibility on social platforms, establishing the basis for feature- and text-driven fake news detection.
Cover for “Liar, Liar Pants on Fire”: A New Benchmark Dataset for Fake News Detection

Abstract

Automatic fake news detection is a challenging problem in deception detection, and it has tremendous real-world political and social impacts. However, statistical approaches to combating fake news has been dramatically limited by the lack of labeled benchmark datasets. In this paper, we present liar: a new, publicly available dataset for fake news detection. We collected a decade-long, 12.8K manually labeled short statements in various contexts from this http URL, which provides detailed analysis report and links to source documents for each case. This dataset can be used for fact-checking research as well. Notably, this new dataset is an order of magnitude larger than previously largest public fake news datasets of similar type. Empirically, we investigate automatic fake news detection based on surface-level linguistic patterns. We have designed a novel, hybrid convolutional neural network to integrate meta-data with text. We show that this hybrid approach can improve a text-only deep learning model.

Table of Contents

  • 1 Introduction
  • 2 liar: a New Benchmark Dataset
  • 3 Automatic Fake News Detection
  • 4 liar: Benchmark Evaluation
  • 4.1 Experimental Settings
  • 4.2 Results
  • 5 Conclusion
  • References

Knowls

  1. Knowl 1 — LIAR Benchmark Dataset for Fake News Detection

    definition

    The LIAR dataset is a publicly available benchmark for automatic fake news detection and fact-checking, comprising 12,836 human-labeled short statements extracted from PolitiFact.com via its API across a decade of political statements (primarily 2007–2016).

    Key characteristics of the dataset include:

    • Truthfulness Labels: Each statement is labeled with one of six fine-grained truthfulness ratings: pants-fire, false, barely-true, half-true, mostly-true, and true. Duplicate PolitiFact verdict categories full-flop, half-flip, and no-flip are merged into false, half-true, and true, respectively.
    • Data Splits: The corpus is split into 10,269 training instances, 1,284 validation instances, and 1,283 test instances. The average statement length is 17.9 tokens.
    • Class Distribution: The classes are relatively balanced; the extreme class pants-fire contains 1,050 instances, while each of the remaining five classes contains between 2,063 and 2,638 instances.
    • Speaker Demographics: The statements originate from diverse sources, with the top affiliations being Republicans (5,687), Democrats (4,150), and unaffiliated entities or social media posts (2,185).
    • Grounding and Quality: Each statement includes a detailed editorial analysis report and links to source documents. Second-stage validation on a random sample of 200 analysis reports showed an inter-annotator agreement of κ=0.82\kappa = 0.82 (Cohen's kappa).
    • Metadata Fields: Each record is annotated with rich metadata, including speaker identity, speaker's current job, home state, party affiliation, context/venue (e.g., campaign speeches, TV/radio interviews, TV ads, news releases, debates, tweets, Facebook posts), subject matter (top subjects include economy, healthcare, taxes, federal-budget, education, jobs, state-budget, candidates-biography, elections, and immigration), and historical speaker credit history.
  2. Knowl 2 — Hybrid Convolutional Neural Network Architecture for Text and Metadata Integration

    model/method

    The Hybrid Convolutional Neural Network (Hybrid CNN) integrates short statement text with categorical and historical metadata to perform fine-grained (6-way) fake news classification.

    The model architecture processes text and metadata through two separate branches before fusing them:

    1. Text Processing Branch: Word tokens of the statement are initialized with 300-dimensional pre-trained word embeddings. A 1D convolutional layer with multiple filter sizes extracts local nn-gram feature maps, followed by a max-over-time pooling operation that yields a fixed-dimensional text representation vector.
    2. Metadata Processing Branch: Metadata attributes (such as speaker, context, party affiliation, current job, state, subject, and credit history) are encoded via a randomly initialized embedding matrix. A convolutional layer captures dependencies across the metadata embeddings, followed by max-pooling on the latent space and a bidirectional Long Short-Term Memory (Bi-LSTM) network layer to model relational representations across metadata fields.
    3. Multimodal Fusion and Output: The max-pooled text representation is concatenated with the final metadata vector output by the Bi-LSTM layer into a joint vector z\mathbf{z}. This vector is passed to a fully connected layer with a softmax activation function to compute class probabilities:
    y^=softmax⁡(Wz+b)\hat{\mathbf{y}} = \operatorname{softmax}(\mathbf{W}\mathbf{z} + \mathbf{b})

    where W\mathbf{W} and b\mathbf{b} are the weight matrix and bias vector of the output layer, respectively.

  3. Knowl 3 — Speaker Credit History Vector Representation and Adjustment

    model/method

    In the LIAR dataset, speaker truthfulness history is represented as a 5-dimensional credit history count vector h∈N5\mathbf{h} \in \mathbb{N}^5:

    h=[cpants-fire,cfalse,cbarely-true,chalf-true,cmostly-true]\mathbf{h} = [c_{\text{pants-fire}}, c_{\text{false}}, c_{\text{barely-true}}, c_{\text{half-true}}, c_{\text{mostly-true}}]

    where each entry ckc_k counts the total number of historical statements by that speaker evaluated under label k∈{pants-fire,false,barely-true,half-true,mostly-true}k \in \{\text{pants-fire}, \text{false}, \text{barely-true}, \text{half-true}, \text{mostly-true}\}.

    To prevent label leakage during training and evaluation, the ground-truth rating of the target statement under prediction must be subtracted from h\mathbf{h} prior to providing the vector to any machine learning model.

  4. Knowl 4 — Classification Accuracy on the LIAR Benchmark Dataset

    data/table

    The empirical performance of text-only classifiers and multimodal hybrid CNN configurations evaluated on the 6-way fine-grained fake news classification task of the LIAR dataset:

    Models Validation Accuracy Test Accuracy
    Text-Only Models
    Majority Baseline 0.204 0.208
    SVMs 0.258 0.255
    Logistic Regression 0.257 0.247
    Bi-LSTMs 0.223 0.233
    CNNs 0.260 0.270
    Hybrid CNN Models (Text + Metadata)
    Text + Subject 0.263 0.235
    Text + Speaker 0.277 0.248
    Text + Job 0.270 0.258
    Text + State 0.246 0.256
    Text + Party 0.259 0.248
    Text + Context 0.251 0.243
    Text + History 0.246 0.241
    Text + All Metadata 0.247 0.274

    The table reports 6-way classification accuracy on the LIAR validation set (1,284 instances) and test set (1,283 instances). Among text-only models, the CNN achieves the highest test accuracy of 0.270. Among hybrid architectures, integrating all available metadata attributes with text (Text + All Metadata) achieves the best overall test accuracy of 0.274.

  5. Knowl 5 — Empirical Comparison of Linguistic Classifiers and Metadata Fusion

    empirical result

    Evaluation of machine learning models on the 6-way LIAR fake news classification task demonstrates three primary findings:

    1. Text CNN Performance: A text-only Convolutional Neural Network (CNN) warm-started with 300-dimensional Google News Word2Vec embeddings achieved a test accuracy of 0.2700.270, significantly outperforming Support Vector Machines (0.2550.255) as determined by a two-tailed paired tt-test (p<0.0001p < 0.0001).
    2. Bi-LSTM Overfitting: A bidirectional LSTM (Bi-LSTM) text model achieved 0.2330.233 test accuracy (0.2230.223 validation accuracy), performing worse than standard linear models (Logistic Regression at 0.2470.247, SVM at 0.2550.255) and the CNN due to overfitting on the short statement texts (average length 17.9 tokens).
    3. Multimodal Metadata Gains: Combining text with all available metadata channels (subject, speaker, job, state, party, context, and credit history) in the Hybrid CNN model achieved the highest test accuracy of 0.2740.274, outperforming all text-only baselines and individual metadata combinations on the test set.
  6. Knowl 6 — Experimental Setup for LIAR Benchmark Evaluation

    experimental setup

    The 6-way fake news classification benchmark on the LIAR dataset evaluates models predicting one of six truthfulness classes (pants-fire, false, barely-true, half-true, mostly-true, true) from statement text and optional metadata.

    • Baselines:
      • Majority Baseline: Always predicts the most frequent label in the training set.
      • Support Vector Machines (SVMs): Linear multi-class SVM implemented using the LibShortText toolkit.
      • Logistic Regression (LR): Regularized logistic regression classifier implemented using the LibShortText toolkit.
      • Bi-LSTM: Bidirectional LSTM implemented in TensorFlow using pre-trained 300-dimensional Google News Word2Vec word embeddings.
      • Text CNN: 1D sentence CNN implemented in TensorFlow using pre-trained 300-dimensional Google News Word2Vec embeddings.
    • Hyperparameters and Training Details:
      • Text CNN: Convolutional filter sizes (2,3,4)(2, 3, 4) with 128 filters per size; dropout keep probability 0.80.8; no L2L_2 regularization penalty; batch size 64; trained with stochastic gradient descent (SGD) for 10 epochs.
      • Hybrid CNN: Convolutional filter sizes 3 and 8 with 10 filters per size; dropout keep probabilities evaluated at 0.50.5 and 0.80.8; trained with SGD for 5 epochs with batch size 64.
      • Hyperparameter Selection: Grid search and validation tuning were conducted strictly on the validation set.
    • Metric: Multi-class classification accuracy on the 1,283 held-out test statements (which is equivalent to macro/micro F-measure due to the balanced class distribution).

Coverage note — None was omitted; all contributed material from this short paper (the LIAR dataset, the Hybrid CNN architecture, experimental configurations, and empirical benchmark results) has been captured.

References

  1. 1.Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. 2011. Natural language processing (almost) from scratch. Journal of Machine Learning Research 12(Aug):2493–2537.
  2. 2.Koby Crammer and Yoram Singer. 2001. On the algorithmic implementation of multiclass kernel-based vector machines. Journal of machine learning research 2(Dec):265–292.
  3. 3.Song Feng, Ritwik Banerjee, and Yejin Choi. 2012. Syntactic stylometry for deception detection. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Short Papers-Volume 2. Association for Computational Linguistics, pages 171–175.
  4. 4.William Ferreira and Andreas Vlachos. 2016. Emergent: a novel data-set for stance classification. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. ACL.
  5. 5.Alex Graves and Jürgen Schmidhuber. 2005. Framewise phoneme classification with bidirectional lstm and other neural network architectures. Neural Networks 18(5):602–610.
  6. 6.Zhen Hai, Peilin Zhao, Peng Cheng, Peng Yang, Xiao-Li Li, Guangxia Li, and Ant Financial. 2016. Deceptive review spam detection via exploiting task relatedness and unlabeled data. In EMNLP.
  7. 7.Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory. Neural computation 9(8):1735–1780.
  8. 8.Cecilia Kang and Adam Goldman. 2016. In washington pizzeria attack, fake news brought real guns. In the New York Times.
  9. 9.Yoon Kim. 2014. Convolutional neural networks for sentence classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP).
  10. 10.Rada Mihalcea and Carlo Strapparava. 2009. The lie detector: Explorations in the automatic recognition of deceptive language. In Proceedings of the ACL-IJCNLP 2009 Conference Short Papers.
  11. 11.Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781 .
  12. 12.Myle Ott, Yejin Choi, Claire Cardie, and Jeffrey T Hancock. 2011. Finding deceptive opinion spam by any stretch of the imagination. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume 1. Association for Computational Linguistics, pages 309–319.
  13. 13.Verónica Pérez-Rosas and Rada Mihalcea. 2015. Experiments in open domain deception detection. In EMNLP. pages 1120–1125.
  14. 14.Andreas Vlachos and Sebastian Riedel. 2014. Fact checking: Task definition and dataset construction. Proceedings of the ACL 2014 Workshop on Language Technology and Computational Social Science .
  15. 15.William Yang Wang and Diyi Yang. 2015. That’s so annoying!!!: A lexical and frame-semantic embedding based data augmentation approach to automatic categorization of annoying behaviors using #petpeeve tweets. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP 2015). ACL, Lisbon, Portugal.
  16. 16.Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Advances in neural information processing systems. pages 649–657.

Citation

MLA
Wang, W. Y. “"Liar, Liar Pants on Fire": A New Benchmark Dataset for Fake News Detection”. arXiv, 2017, http://arxiv.org/abs/1705.00648v1.
APA
Wang, W. Y. (2017). "Liar, Liar Pants on Fire": A New Benchmark Dataset for Fake News Detection. arXiv. http://arxiv.org/abs/1705.00648v1
Chicago
Wang, W. Y. 2017. “"Liar, Liar Pants on Fire": A New Benchmark Dataset for Fake News Detection”. arXiv. http://arxiv.org/abs/1705.00648v1.
Harvard
Wang, W.Y. (2017) “"Liar, Liar Pants on Fire": A New Benchmark Dataset for Fake News Detection”, arXiv [Preprint]. Available at: http://arxiv.org/abs/1705.00648v1.
Vancouver
1. Wang WY (2017) "Liar, Liar Pants on Fire": A New Benchmark Dataset for Fake News Detection. arXiv

BibTeX

@article{wang2017liar,
  title = {"Liar, Liar Pants on Fire": A New Benchmark Dataset for Fake News Detection},
  author = {Wang, William Yang},
  year = {2017},
  journal = {arXiv},
  url = {http://arxiv.org/abs/1705.00648v1},
  eprint = {1705.00648}
}
Metadata:arXiv

Access the Paper

This paper is available from its original source. Click below to access the PDF.

Open PDF
License: https://creativecommons.org/licenses/by/4.0/