"Liar, Liar Pants on Fire": A New Benchmark Dataset for Fake News Detection

William Yang Wang
Department of Computer Science
University of California, Santa Barbara
Santa Barbara, CA 93106 USA
Abstract

Automatic fake news detection is a challenging problem in deception detection, and it has tremendous real-world political and social impacts. However, statistical approaches to combating fake news has been dramatically limited by the lack of labeled benchmark datasets. In this paper, we present LIAR: a new, publicly available dataset for fake news detection. We collected a decade-long, 12.8K manually labeled short statements in various contexts from POLITIFACT.COM, which provides detailed analysis report and links to source documents for each case. This dataset can be used for fact-checking research as well. Notably, this new dataset is an order of magnitude larger than previously largest public fake news datasets of similar type. Empirically, we investigate automatic fake news detection based on surface-level linguistic patterns. We have designed a novel, hybrid convolutional neural network to integrate metadata with text. We show that this hybrid approach can improve a text-only deep learning model.

1. Introduction

In this past election cycle for the 45th President of the United States, the world has witnessed a growing epidemic of fake news. The plague of fake news not only poses serious threats to the integrity of journalism, but has also created turmoils in the political world. The worst real-world impact is that fake news seems to create real-life fears: last year, a man carried an AR-15 rifle and walked in a Washington DC Pizzeria, because he recently read online that "this pizzeria was harboring young children as sex slaves as part of a child-abuse ring led by Hillary Clinton". The man was later arrested by police, and he was charged for firing an assault rifle in the restaurant (Kang and Goldman, 2016).

The broadly-related problem of deception detection (Mihalcea and Strapparava, 2009) is not new to the natural language processing community. A relatively early study by Ott et al. (2011) focuses on detecting deceptive review opinions in sentiment analysis, using a crowdsourcing approach to create training data for the positive class, and then combine with truthful opinions from TripAdvisor. Recent studies have also proposed stylometric (Feng et al., 2012), semi-supervised learning (Hai et al., 2016), and linguistic approaches (Perez-Rosas and Mihalcea, 2015) to detect deceptive text on crowdsourced datasets. Even though crowdsourcing is an important approach to create labeled training data, there is a mismatch between training and testing. When testing on real-world review datasets, the results could be suboptimal since the positive training data was created in a completely different, simulated platform.

The problem of fake news detection is more challenging than detecting deceptive reviews, since the political language on TV interviews, posts on Facebook and Twitters are mostly short statements. However, the lack of manually labeled fake news dataset is still a bottleneck for advancing computational-intensive, broad-coverage models in this direction. Vlachos and Riedel (2014) are the first to release a public fake news detection and fact-checking dataset, but it only includes 221 statements, which does not permit machine learning based assessments.

To address these issues, we introduce the LIAR dataset, which includes 12,836 short statements labeled for truthfulness, subject, context/venue, speaker, state, party, and prior history. With such volume and a time span of a decade, LIAR is an order of magnitude larger than the currently available resources (Vlachos and Riedel, 2014; Ferreira and Vlachos, 2016) of similar type. Additionally, in contrast to crowdsourced datasets, the instances in LIAR are collected in a grounded, more natural context, such as political debate, TV ads, Facebook posts, tweets, interview, news release, etc. In each case, the labeler provides a lengthy analysis report to ground each judgment, and the links to all supporting documents are also provided.

Empirically, we have evaluated several popular learning based methods on this dataset. The baselines include logistic regression, support vector machines, long short-term memory networks (Hochreiter and Schmidhuber, 1997), and a convolutional neural network model (Kim, 2014). We further introduce a neural network architecture to integrate text and meta-data. Our experiment suggests that this approach improves the performance of a strong text-only convolutional neural networks baseline.

2. LIAR: a New Benchmark Dataset

The major resources for deceptive detection of reviews are crowdsourced datasets (Ott et al., 2011; Perez-Rosas and Mihalcea, 2015). They are very useful datasets to study deception detection, but the positive training data are collected from a simulated environment. More importantly, these datasets are not suitable for fake statements detection, since the fake news on TVs and social media are much shorter than customer reviews.

Vlachos and Riedel (2014) are the first to construct fake news and fact-checking datasets. They obtained 221 statements from CHANNEL 4 and POLITIFACT.COM, a Pulitzer Prize-winning website. In particular, PolitiFact covers a wide-range of political topics, and they provide detailed judgments with fine-grained labels. Recently, Ferreira and Vlachos (2016) have released the Emergent dataset, which includes 300 labeled rumors from PolitiFact. However, with less than a thousand samples, it is impractical to use these datasets as benchmarks for developing and evaluating machine learning algorithms for fake news detection.

StatementSpeakerContextLabelJustification
"The last quarter, it was just announced, our gross domestic product was below zero. Who ever heard of this? It's never below zero."Donald Trumppresidential announcement speechPants on FireAccording to Bureau of Economic Analysis and National Bureau of Economic Research, the growth in the gross domestic product has been below zero 42 times over 68 years. That's a lot more than "never." We rate his claim Pants on Fire!
"Newly Elected Republican Senators Sign Pledge to Eliminate Food Stamp Program in 2015."Facebook postssocial media postingPants on FireMore than 115,000 social media users passed along a story headlined, "Newly Elected Republican Senators Sign Pledge to Eliminate Food Stamp Program in 2015." But they failed to do due diligence and were snookered, since the story came from a publication that bills itself (quietly) as a "satirical, parody website." We rate the claim Pants on Fire.
"Under the health care law, everybody will have lower rates, better quality care and better access."Nancy Pelosion 'Meet the Press'FalseEven the study that Pelosi's staff cited as the source of that statement suggested that some people would pay more for health insurance. Analysis at the state level found the same thing. The general understanding of the word "everybody" is every person. The predictions don't back that up. We rule this statement False.
Table 1: Some random excerpts from the LIAR dataset.

Therefore, it is of crucial significance to introduce a larger dataset to facilitate the development of computational approaches to fake news detection and automatic fact-checking.

We show some random snippets from our dataset in Figure 1. The LIAR dataset includes 12.8K human labeled short statements from POLITIFACT.COM's API, and each statement is evaluated by a POLITIFACT.COM editor for its truthfulness. After initial analysis, we found duplicate labels, and merged the full-flop, half-flip, no-flip labels into false, half-true, true labels respectively. We consider six fine-grained labels for the truthfulness ratings: pants-fire, false, barely-true, half-true, mostly-true, and true. The distribution of labels in the LIAR dataset is relatively well-balanced: except for 1,050 pants-fire cases, the instances for all other labels range from 2,063 to 2,638. We randomly sampled 200 instances to examine the accompanied lengthy analysis reports and rulings. Not that fact-checking is not a classic labeling task in NLP. The verdict requires extensive training in journalism for finding relevant evidence. Therefore, for second-stage verifications, we went through a randomly sampled subset of the analysis reports to check if we agreed with the reporters' analysis. The agreement rate measured by Cohens kappa was 0.82. We show the corpus statistics in Table 1. The statement dates are primarily from 2007-2016.

Dataset Statistics
Training set size10,269
Validation set size1,284
Testing set size1,283
Avg. statement length (tokens)17.9
Top-3 Speaker Affiliations
  Democrats4,150
  Republicans5,687
  None (e.g., FB posts)2,185
Table 1: The LIAR dataset statistics.

The speakers in the LIAR dataset include a mix of democrats and republicans, as well as a significant amount of posts from online social media. We include a rich set of meta-data for each speaker—in addition to party affiliations, current job, home state, and credit history are also provided. In particular, the credit history includes the historical counts of inaccurate statements for each speaker. For example, Mitt Romney has a credit history vector h = {19, 32, 34, 58, 33}, which corresponds to his counts of "pants on fire", "false", "barely true", "half true", "mostly true" for historical statements. Since this vector also includes the count for the current statement, it is important to subtract the current label from the credit history when using this meta data vector in prediction experiments.

These statements are sampled from various of contexts/venues, and the top categories include news releases, TV/radio interviews, campaign speeches, TV ads, tweets, debates, Facebook posts, etc. To ensure a broad coverage of the topics, there is also a diverse set of subjects discussed by the speakers. The top-10 most discussed subjects in the dataset are economy, healthcare, taxes, federal-budget, education, jobs, state-budget, candidates-biography, elections, and immigration.

3. Automatic Fake News Detection

One of the most obvious applications of our dataset is to facilitate the development of machine learning models for automatic fake news detection. In this task, we frame this as a 6-way multi-class text classification problem. And the research questions are:

Since convolutional neural networks architectures (CNNs) (Collobert et al., 2011; Kim, 2014; Zhang et al., 2015) have obtained the state-of-the-art results on many text classification datasets, we build our neural networks model based on a recently proposed CNN model (Kim, 2014).

The proposed hybrid Convolutional Neural Networks framework for integrating text and meta-data
Figure 2: The proposed hybrid Convolutional Neural Networks framework for integrating text and meta-data.

Figure 2 shows the overview of our hybrid convolutional neural network for integrating text and meta-data.

We randomly initialize a matrix of embedding vectors to encode the metadata embeddings. We use a convolutional layer to capture the dependency between the meta-data vector(s). Then, a standard max-pooling operation is performed on the latent space, followed by a bi-directional LSTM layer. We then concatenate the max-pooled text representations with the meta-data representation from the bi-directional LSTM, and feed them to fully connected layer with a softmax activation function to generate the final prediction.

4. LIAR: Benchmark Evaluation

In this section, we first describe the experimental setup, and the baselines. Then, we present the empirical results and compare various models.

4.1 Experimental Settings

We used five baselines: a majority baseline, a regularized logistic regression classifier (LR), a support vector machine classifier (SVM) (Crammer and Singer, 2001), a bi-directional long short-term memory networks model (Bi-LSTMs) (Hochreiter and Schmidhuber, 1997; Graves and Schmidhuber, 2005), and a convolutional neural network model (CNNs) (Kim, 2014). For LR and SVM, we used the LIBSHORTTEXT toolkit, which was shown to provide very strong performances on short text classification problems (Wang and Yang, 2015). For Bi-LSTMs and CNNs, we used TensorFlow for the implementation. We used pre-trained 300-dimensional word2vec embeddings from Google News (Mikolov et al., 2013) to warm-start the text embeddings. We strictly tuned all the hyperparameters on the validation dataset.

The best filter sizes for the CNN model was (2,3,4). In all cases, each size has 128 filters. The dropout keep probabilities was optimized to 0.8, while no L2 penalty was imposed. The batch size for stochastic gradient descent optimization was set to 64, and the learning process involves 10 passes over the training data for text model. For the hybrid model, we use 3 and 8 as filter sizes, and the number of filters was set to 10. We considered 0.5 and 0.8 as dropout probabilities. The hybrid model requires 5 training epochs.

We used grid search to tune the hyperparameters for LR and SVM models. We chose accuracy as the evaluation metric, since we found that the accuracy results from various models were equivalent to f-measures on this balanced dataset.

4.2 Results

We outline our empirical results in Table 2. First, we compare various models using text features only. We see that the majority baseline on this dataset gives about 0.204 and 0.208 accuracy on the validation and test sets respectively. Standard text classifier such as SVMs and LR models obtained significant improvements. Due to overfitting, the Bi-LSTMs did not perform well. The CNNs outperformed all models, resulting in an accuracy of 0.270 on the heldout test set. We compare the predictions from the CNN model with SVMs via a two-tailed paired t-test, and CNN was significantly better (p < .0001). When considering all meta-data and text, the model achieved the best result on the test data.

ModelsValid.Test
Majority0.2040.208
SVMs0.2580.255
Logistic Regression0.2570.247
Bi-LSTMs0.2230.233
CNNs0.2600.270
Hybrid CNNs
  Text + Subject0.2630.235
  Text + Speaker0.2770.248
  Text + Job0.2700.258
  Text + State0.2460.256
  Text + Party0.2590.248
  Text + Context0.2510.243
  Text + History0.2460.241
  Text + All0.2470.274
Table 2: The evaluation results on the LIAR dataset. The top section: text-only models. The bottom: text + meta-data hybrid models.

5. Conclusion

We introduced LIAR, a new dataset for automatic fake news detection. Compared to prior datasets, LIAR is an order of a magnitude larger, which enables the development of statistical and computational approaches to fake news detection. LIAR's authentic, real-world short statements from various contexts with diverse speakers also make the research on developing broad-coverage fake news detector possible. We show that when combining meta-data with text, significant improvements can be achieved for fine-grained fake news detection. Given the detailed analysis report and links to source documents in this dataset, it is also possible to explore the task of automatic fact-checking over knowledge base in the future. Our corpus can also be used for stance classification, argument mining, topic modeling, rumor detection, and political NLP research.

6. References

Ronan Collobert, Jason Weston, Leon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. 2011. Natural language processing (almost) from scratch. Journal of Machine Learning Research 12(Aug):2493–2537.
Koby Crammer and Yoram Singer. 2001. On the algorithmic implementation of multiclass kernel-based vector machines. Journal of machine learning research 2(Dec):265–292.
Song Feng, Ritwik Banerjee, and Yejin Choi. 2012. Syntactic stylometry for deception detection. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Short Papers-Volume 2. Association for Computational Linguistics, pages 171–175.
William Ferreira and Andreas Vlachos. 2016. Emergent: a novel data-set for stance classification. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. ACL.
Alex Graves and Jurgen Schmidhuber. 2005. Frame-wise phoneme classification with bidirectional lstm and other neural network architectures. Neural Networks 18(5):602–610.
Zhen Hai, Peilin Zhao, Peng Cheng, Peng Yang, Xiao-Li Li, Guangxia Li, and Ant Financial. 2016. Deceptive review spam detection via exploiting task relatedness and unlabeled data. In EMNLP.
Sepp Hochreiter and Jurgen Schmidhuber. 1997. Long short-term memory. Neural computation 9(8):1735–1780.
Cecilia Kang and Adam Goldman. 2016. In washington pizzeria attack, fake news brought real guns. In the New York Times.
Yoon Kim. 2014. Convolutional neural networks for sentence classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP).
Rada Mihalcea and Carlo Strapparava. 2009. The lie detector: Explorations in the automatic recognition of deceptive language. In Proceedings of the ACL-IJCNLP 2009 Conference Short Papers.
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.
Myle Ott, Yejin Choi, Claire Cardie, and Jeffrey T Hancock. 2011. Finding deceptive opinion spam by any stretch of the imagination. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume 1. Association for Computational Linguistics, pages 309–319.
Veronica Perez-Rosas and Rada Mihalcea. 2015. Experiments in open domain deception detection. In EMNLP. pages 1120–1125.
Andreas Vlachos and Sebastian Riedel. 2014. Fact checking: Task definition and dataset construction. Proceedings of the ACL 2014 Workshop on Language Technology and Computational Social Science.
William Yang Wang and Diyi Yang. 2015. That's so annoying!!!: A lexical and frame-semantic embedding based data augmentation approach to automatic categorization of annoying behaviors using #petpeeve tweets. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP 2015). ACL, Lisbon, Portugal.
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Advances in neural information processing systems. pages 649–657.

Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, pages 422-426

"Лжец, лжец, штаны в огне": Новый эталонный набор данных для обнаружения фейковых новостей

Уильям Янг Ванг
Факультет компьютерных наук
Калифорнийский университет в Санта-Барбаре
Санта-Барбара, Калифорния 93106 США
Аннотация

Автоматическое обнаружение фейковых новостей является сложной проблемой в обнаружении обмана и оказывает огромное влияние на реальную политическую и социальную жизнь. Однако статистические подходы к борьбе с фейковыми новостями были значительно ограничены отсутствием размеченных эталонных наборов данных. В этой статье мы представляем LIAR: новый общедоступный набор данных для обнаружения фейковых новостей. Мы собрали десятилетние, 12,8 тысяч вручную размеченных коротких утверждений в различных контекстах с POLITIFACT.COM, который предоставляет подробный аналитический отчет и ссылки на исходные документы по каждому случаю. Этот набор данных также может использоваться для исследований по проверке фактов. Примечательно, что этот новый набор данных на порядок больше, чем предыдущие крупнейшие общедоступные наборы данных о фейковых новостях подобного типа. Эмпирически мы исследуем автоматическое обнаружение фейковых новостей на основе поверхностных лингвистических паттернов. Мы разработали новую гибридную сверточную нейронную сеть для интеграции метаданных с текстом. Мы показываем, что этот гибридный подход может улучшить модель глубокого обучения, работающую только с текстом.

1. Введение

В прошлом избирательном цикле для 45-го президента США мир стал свидетелем растущей эпидемии фейковых новостей. Чума фейковых новостей не только представляет серьезную угрозу целостности журналистики, но и создала потрясения в политическом мире. Худшим реальным последствием является то, что фейковые новости, по-видимому, создают реальные страхи: в прошлом году мужчина с винтовкой AR-15 вошел в пиццерию в Вашингтоне, округ Колумбия, потому что недавно прочитал в интернете, что "эта пиццерия укрывает маленьких детей в качестве секс-рабов в рамках кольца жестокого обращения с детьми под руководством Хиллари Клинтон". Мужчина был позже арестован полицией, и ему было предъявлено обвинение в стрельбе из штурмовой винтовки в ресторане (Kang and Goldman, 2016).

Широко связанная проблема обнаружения обмана (Mihalcea and Strapparava, 2009) не нова для сообщества обработки естественного языка. Относительно раннее исследование Ott et al. (2011) фокусируется на обнаружении обманчивых обзоров в анализе тональности, используя краудсорсинговый подход для создания обучающих данных для положительного класса, а затем комбинируя их с правдивыми мнениями из TripAdvisor. Недавние исследования также предложили стилометрические (Feng et al., 2012), полуконтролируемое обучение (Hai et al., 2016) и лингвистические подходы (Perez-Rosas and Mihalcea, 2015) для обнаружения обманчивого текста на краудсорсинговых наборах данных. Хотя краудсорсинг является важным подходом для создания размеченных обучающих данных, существует несоответствие между обучением и тестированием. При тестировании на реальных наборах данных отзывов результаты могут быть неоптимальными, поскольку положительные обучающие данные были созданы на совершенно другой, смоделированной платформе.

Проблема обнаружения фейковых новостей является более сложной, чем обнаружение обманчивых отзывов, поскольку политический язык в телевизионных интервью, постах на Facebook и Twitter в основном состоит из коротких утверждений. Однако отсутствие вручную размеченного набора данных о фейковых новостях по-прежнему является узким местом для продвижения вычислительно интенсивных моделей с широким охватом в этом направлении. Vlachos и Riedel (2014) первыми выпустили общедоступный набор данных для обнаружения фейковых новостей и проверки фактов, но он включает только 221 утверждение, что не позволяет проводить оценки на основе машинного обучения.

Для решения этих проблем мы представляем набор данных LIAR, который включает 12 836 коротких утверждений, размеченных по правдивости, теме, контексту/месту, спикеру, штату, партии и предыдущей истории. При таком объеме и временном промежутке в десятилетие LIAR на порядок больше, чем доступные в настоящее время ресурсы (Vlachos and Riedel, 2014; Ferreira and Vlachos, 2016) аналогичного типа. Кроме того, в отличие от краудсорсинговых наборов данных, экземпляры в LIAR собираются в обоснованном, более естественном контексте, таком как политические дебаты, телевизионная реклама, посты в Facebook, твиты, интервью, пресс-релизы и т.д. В каждом случае разметчик предоставляет подробный аналитический отчет, обосновывающий каждое суждение, а также ссылки на все подтверждающие документы.

Эмпирически мы оценили несколько популярных методов на основе обучения на этом наборе данных. Базовые модели включают логистическую регрессию, метод опорных векторов, сети долгой краткосрочной памяти (Hochreiter and Schmidhuber, 1997) и модель сверточной нейронной сети (Kim, 2014). Мы также представляем архитектуру нейронной сети для интеграции текста и метаданных. Наш эксперимент показывает, что этот подход улучшает производительность сильной базовой модели сверточных нейронных сетей, работающей только с текстом.

2. LIAR: новый эталонный набор данных

Основными ресурсами для обнаружения обмана в отзывах являются краудсорсинговые наборы данных (Ott et al., 2011; Perez-Rosas and Mihalcea, 2015). Это очень полезные наборы данных для изучения обнаружения обмана, но положительные обучающие данные собираются в смоделированной среде. Что более важно, эти наборы данных не подходят для обнаружения фейковых утверждений, поскольку фейковые новости на телевидении и в социальных сетях намного короче, чем отзывы клиентов.

Vlachos и Riedel (2014) первыми создали наборы данных для обнаружения фейковых новостей и проверки фактов. Они получили 221 утверждение от CHANNEL 4 и POLITIFACT.COM, веб-сайта, удостоенного Пулитцеровской премии. В частности, PolitiFact охватывает широкий спектр политических тем и предоставляет подробные суждения с детальными метками. Недавно Ferreira и Vlachos (2016) выпустили набор данных Emergent, который включает 300 размеченных слухов из PolitiFact. Однако при менее чем тысяче образцов непрактично использовать эти наборы данных в качестве эталонов для разработки и оценки алгоритмов машинного обучения для обнаружения фейковых новостей.

УтверждениеСпикерКонтекстМеткаОбоснование
"The last quarter, it was just announced, our gross domestic product was below zero. Who ever heard of this? It's never below zero."Дональд Трамппрезентационная речь на президентских выборахPants on FireСогласно Бюро экономического анализа и Национальному бюро экономических исследований, рост валового внутреннего продукта был ниже нуля 42 раза за 68 лет. Это намного больше, чем «никогда». Мы оцениваем это заявление как Pants on Fire!
"Newly Elected Republican Senators Sign Pledge to Eliminate Food Stamp Program in 2015."Посты в Facebookпубликация в социальных сетяхPants on FireБолее 115 000 пользователей социальных сетей распространили историю под заголовком «Вновь избранные республиканские сенаторы подписывают обязательство об отмене программы продовольственных талонов в 2015 году». Но они не проявили должной осмотрительности и были обмануты, поскольку история поступила из издания, которое (тихо) позиционирует себя как «сатирический, пародийный веб-сайт». Мы оцениваем это заявление как Pants on Fire.
"Under the health care law, everybody will have lower rates, better quality care and better access."Нэнси Пелосина передаче 'Meet the Press'FalseДаже исследование, на которое ссылался персонал Пелоси как на источник этого утверждения, предполагало, что некоторые люди будут платить больше за медицинскую страховку. Анализ на уровне штатов обнаружил то же самое. Общее понимание слова «все» - каждый человек. Прогнозы этого не подтверждают. Мы считаем это утверждение ложным.
Таблица 1: Некоторые случайные выдержки из набора данных LIAR.

Поэтому имеет решающее значение введение более крупного набора данных для облегчения развития вычислительных подходов к обнаружению фейковых новостей и автоматической проверке фактов.

Мы показываем некоторые случайные фрагменты из нашего набора данных на Рисунке 1. Набор данных LIAR включает 12,8 тысяч вручную размеченных коротких утверждений из API POLITIFACT.COM, и каждое утверждение оценивается редактором POLITIFACT.COM на предмет его правдивости. После первоначального анализа мы обнаружили дублирующиеся метки и объединили метки full-flop, half-flip, no-flip в метки false, half-true, true соответственно. Мы рассматриваем шесть детальных меток для оценок правдивости: pants-fire, false, barely-true, half-true, mostly-true и true. Распределение меток в наборе данных LIAR относительно хорошо сбалансировано: за исключением 1050 случаев pants-fire, экземпляры для всех других меток варьируются от 2063 до 2638. Мы случайным образом отобрали 200 экземпляров, чтобы измотреть сопровождающие подробные аналитические отчеты и решения. Отметим, что проверка фактов — это не классическая задача разметки в NLP. Вердикт требует обширной подготовки в журналистике для поиска соответствующих доказательств. Поэтому для проверок второго этапа мы просмотрели случайно выбранное подмножество аналитических отчетов, чтобы проверить, согласны ли мы с анализом репортеров. Коэффициент согласия, измеренный каппой Коэна, составил 0,82. Мы показываем статистику корпуса в Таблице 1. Даты утверждений в основном с 2007 по 2016 год.

Статистика набора данных
Размер обучающего набора10,269
Размер валидационного набора1,284
Размер тестового набора1,283
Средняя длина утверждения (токены)17.9
Топ-3 принадлежностей спикеров
  Демократы4,150
  Республиканцы5,687
  Нет (например, посты FB)2,185
Таблица 1: Статистика набора данных LIAR.

Спикеры в наборе данных LIAR включают смесь демократов и республиканцев, а также значительное количество постов из онлайн-социальных сетей. Мы включаем богатый набор метаданных для каждого спикера — помимо партийной принадлежности, также предоставляются текущая работа, штат проживания и кредитная история. В частности, кредитная история включает историческое количество неточных утверждений для каждого спикера. Например, Митт Ромни имеет вектор кредитной истории h = {19, 32, 34, 58, 33}, который соответствует его количеству "pants on fire", "false", "barely true", "half true", "mostly true" для исторических утверждений. Поскольку этот вектор также включает количество для текущего утверждения, важно вычесть текущую метку из кредитной истории при использовании этого вектора метаданных в экспериментах по прогнозированию.

Эти утверждения взяты из различных контекстов/мест, и основные категории включают пресс-релизы, теле/радиоинтервью, предвыборные речи, телевизионную рекламу, твиты, дебаты, посты в Facebook и т.д. Чтобы обеспечить широкий охват тем, также существует разнообразный набор тем, обсуждаемых спикерами. Топ-10 наиболее обсуждаемых тем в наборе данных: экономика, здравоохранение, налоги, федеральный бюджет, образование, рабочие места, бюджет штата, биография кандидатов, выборы и иммиграция.

3. Автоматическое обнаружение фейковых новостей

Одним из наиболее очевидных применений нашего набора данных является облегчение разработки моделей машинного обучения для автоматического обнаружения фейковых новостей. В этой задаче мы рассматриваем это как 6-классовую многоклассовую классификацию текста. И исследовательские вопросы:

Поскольку архитектуры сверточных нейронных сетей (CNN) (Collobert et al., 2011; Kim, 2014; Zhang et al., 2015) достигли наилучших результатов на многих наборах данных классификации текста, мы строим нашу модель нейронных сетей на основе недавно предложенной модели CNN (Kim, 2014).

Предлагаемая гибридная структура сверточных нейронных сетей для интеграции текста и метаданных
Рисунок 2: Предлагаемая гибридная структура сверточных нейронных сетей для интеграции текста и метаданных.

Рисунок 2 показывает обзор нашей гибридной сверточной нейронной сети для интеграции текста и метаданных.

Мы случайным образом инициализируем матрицу векторов вложений для кодирования вложений метаданных. Мы используем сверточный слой для захвата зависимости между векторами метаданных. Затем стандартная операция max-pooling выполняется в латентном пространстве, за которым следует двунаправленный слой LSTM. Затем мы объединяем max-pooled текстовые представления с представлением метаданных из двунаправленного LSTM и передаем их полностью связанному слою с функцией активации softmax для генерации окончательного прогноза.

4. LIAR: Оценка эталона

В этом разделе мы сначала описываем экспериментальную установку и базовые модели. Затем мы представляем эмпирические результаты и сравниваем различные модели.

4.1 Экспериментальные настройки

Мы использовали пять базовых моделей: базовую модель большинства, регуляризованный классификатор логистической регрессии (LR), классификатор метода опорных векторов (SVM) (Crammer and Singer, 2001), модель двунаправленных сетей долгой краткосрочной памяти (Bi-LSTMs) (Hochreiter and Schmidhuber, 1997; Graves and Schmidhuber, 2005) и модель сверточной нейронной сети (CNNs) (Kim, 2014). Для LR и SVM мы использовали инструментарий LIBSHORTTEXT, который, как было показано, обеспечивает очень высокую производительность на задачах классификации короткого текста (Wang and Yang, 2015). Для Bi-LSTMs и CNNs мы использовали TensorFlow для реализации. Мы использовали предобученные 300-мерные вложения word2vec из Google News (Mikolov et al., 2013) для "разогрева" текстовых вложений. Мы строго настраивали все гиперпараметры на валидационном наборе данных.

Лучшие размеры фильтров для модели CNN были (2,3,4). Во всех случаях каждый размер имеет 128 фильтров. Вероятность удержания dropout была оптимизирована до 0,8, при этом штраф L2 не налагался. Размер пакета для оптимизации стохастического градиентного спуска был установлен на 64, а процесс обучения включает 10 проходов по обучающим данным для текстовой модели. Для гибридной модели мы используем 3 и 8 в качестве размеров фильтров, а количество фильтров было установлено на 10. Мы рассматривали 0,5 и 0,8 в качестве вероятностей dropout. Гибридная модель требует 5 эпох обучения.

Мы использовали поиск по сетке для настройки гиперпараметров для моделей LR и SVM. Мы выбрали точность в качестве метрики оценки, поскольку обнаружили, что результаты точности различных моделей были эквивалентны f-мерам на этом сбалансированном наборе данных.

4.2 Результаты

Мы излагаем наши эмпирические результаты в Таблице 2. Сначала мы сравниваем различные модели, использующие только текстовые признаки. Мы видим, что базовая модель большинства на этом наборе данных дает точность около 0,204 и 0,208 на валидационном и тестовом наборах соответственно. Стандартные текстовые классификаторы, такие как модели SVM и LR, получили значительные улучшения. Из-за переобучения Bi-LSTMs показали себя не очень хорошо. CNN превзошли все модели, достигнув точности 0,270 на удержанном тестовом наборе. Мы сравниваем прогнозы модели CNN с SVM с помощью двустороннего парного t-теста, и CNN была значительно лучше (p < 0,0001). При рассмотрении всех метаданных и текста модель достигла наилучшего результата на тестовых данных.

МоделиВалид.Тест
Большинство0.2040.208
SVM0.2580.255
Логистическая регрессия0.2570.247
Bi-LSTM0.2230.233
CNN0.2600.270
Гибридные CNN
  Текст + Тема0.2630.235
  Текст + Спикер0.2770.248
  Текст + Работа0.2700.258
  Текст + Штат0.2460.256
  Текст + Партия0.2590.248
  Текст + Контекст0.2510.243
  Текст + История0.2460.241
  Текст + Все0.2470.274
Таблица 2: Результаты оценки на наборе данных LIAR. Верхняя часть: модели только с текстом. Нижняя: гибридные модели текст + метаданные.

5. Заключение

Мы представили LIAR, новый набор данных для автоматического обнаружения фейковых новостей. По сравнению с предыдущими наборами данных, LIAR на порядок больше, что позволяет разрабатывать статистические и вычислительные подходы к обнаружению фейковых новостей. Подлинные, реальные короткие утверждения LIAR из различных контекстов с разнообразными спикерами также делают возможными исследования по разработке детектора фейковых новостей с широким охватом. Мы показываем, что при объединении метаданных с текстом могут быть достигнуты значительные улучшения для детального обнаружения фейковых новостей. Учитывая подробный аналитический отчет и ссылки на исходные документы в этом наборе данных, также возможно исследовать задачу автоматической проверки фактов по базе знаний в будущем. Наш корпус также может использоваться для классификации позиций, извлечения аргументов, тематического моделирования, обнаружения слухов и политических исследований NLP.

6. Список литературы

Ronan Collobert, Jason Weston, Leon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. 2011. Natural language processing (almost) from scratch. Journal of Machine Learning Research 12(Aug):2493–2537.
Koby Crammer and Yoram Singer. 2001. On the algorithmic implementation of multiclass kernel-based vector machines. Journal of machine learning research 2(Dec):265–292.
Song Feng, Ritwik Banerjee, and Yejin Choi. 2012. Syntactic stylometry for deception detection. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Short Papers-Volume 2. Association for Computational Linguistics, pages 171–175.
William Ferreira and Andreas Vlachos. 2016. Emergent: a novel data-set for stance classification. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. ACL.
Alex Graves and Jurgen Schmidhuber. 2005. Frame-wise phoneme classification with bidirectional lstm and other neural network architectures. Neural Networks 18(5):602–610.
Zhen Hai, Peilin Zhao, Peng Cheng, Peng Yang, Xiao-Li Li, Guangxia Li, and Ant Financial. 2016. Deceptive review spam detection via exploiting task relatedness and unlabeled data. In EMNLP.
Sepp Hochreiter and Jurgen Schmidhuber. 1997. Long short-term memory. Neural computation 9(8):1735–1780.
Cecilia Kang and Adam Goldman. 2016. In washington pizzeria attack, fake news brought real guns. In the New York Times.
Yoon Kim. 2014. Convolutional neural networks for sentence classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP).
Rada Mihalcea and Carlo Strapparava. 2009. The lie detector: Explorations in the automatic recognition of deceptive language. In Proceedings of the ACL-IJCNLP 2009 Conference Short Papers.
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.
Myle Ott, Yejin Choi, Claire Cardie, and Jeffrey T Hancock. 2011. Finding deceptive opinion spam by any stretch of the imagination. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume 1. Association for Computational Linguistics, pages 309–319.
Veronica Perez-Rosas and Rada Mihalcea. 2015. Experiments in open domain deception detection. In EMNLP. pages 1120–1125.
Andreas Vlachos and Sebastian Riedel. 2014. Fact checking: Task definition and dataset construction. Proceedings of the ACL 2014 Workshop on Language Technology and Computational Social Science.
William Yang Wang and Diyi Yang. 2015. That's so annoying!!!: A lexical and frame-semantic embedding based data augmentation approach to automatic categorization of annoying behaviors using #petpeeve tweets. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP 2015). ACL, Lisbon, Portugal.
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Advances in neural information processing systems. pages 649–657.

Труды 55-й ежегодной встречи Ассоциации компьютерной лингвистики, страницы 422-426