NoFake at CheckThat! 2021: Fake News Detection Using BERT

Sushma Kumari
Independent Researcher
Abstract

Much research has been done for debunking and analysing fake news. Many researchers study fake news detection in the last year, but many are limited to social media data. Currently, multiples fact-checkers are publishing their results in various formats. Also, multiple fact-checkers use different labels for the fake news, making it difficult to make a generalisable classifier. With the merge classes, the performance of the machine model can be enhanced. This domain categorisation will help group the article, which will help save the manual effort in assigning the claim verification. In this paper, we have presented BERT based classification model to predict the domain and classification. We have also used additional data from fact-checked articles. We have achieved a macro F1 score of 83.76% for Task 3A and 85.55% for Task 3B using the additional training data.

Keywords: Fake News, Multi-class Classification, Domain Classification, Misinformation

1. Introduction

Data journalism is a new journalistic discipline that focuses mainly on data-driven research and presentation formats. However, a fundamental problem of data journalism and classical journalism is that much data of journalistic interest is only available in the unstructured form: as texts, tables and graphics in documents of various types (Word, PDF, e-mail, etc.) or on websites. Also, there is a lack of a centralised data hub for gathering information from different sources.

There is an increasing amount of fake news in the media, social media, and other web sources. Much research has been done for fake news detection and debunking of fake news [1]. In the last two decades, there is a tremendous increase in the spread of misinformation, which is also reflected by the number of fact-checking websites [1]. Fact-checking websites can help to investigate claims and assist citizens in determining whether the information used in an article is true or not. More than 213 fact-checking websites are working in 40+ languages across 100+ countries [2, 3, 4].

There is no standard protocol for fact-checking services across different fact-checkers, and they do not publish their proofed articles in a standard format, which leads to several conflicts. Shahi et al. [3] discuss the need to detect news articles potentially containing fake information. In this paper, we discussed the method used for the fake news detection for the shared task at CheckThat!. The remainder of the paper is organised as related work describes the past work, task description gives an overview of the task, Experiment section emphasises on the experiment details and Conclusion, and Future Work focuses on the conclusion and future aspect of the work.

2. Related Work

Fact-checking is a damage control method that is both essential and not scalable. It might be hard to take the human component out of the picture any time soon. But still, automatic fake news classification could help reduce the workload for the fact-checkers. The fake content is spread in multiple formats; many of them are repurposed by changing the text, location, etc.

Fake news detection is a complex problem. Research has been done on fake news detection using social media data like tweets [2], YouTube videos [3], but less research has focused on the news articles. One of the primary reasons is the lack of corpus of news articles. Fake news is spread in several domains like crime, election, economy.

There is a lack of corpus to train Machine Learning based on fact-checking. In the last few years, a collection of small datasets related to fact-checking has been published. These datasets are a mixture of fake news topics, For instance, US election 2016[5]. Still, fact-checking is dominant in the English language and few sources like Snopes, Politifact etc.

Multi-FC corpus describes the different kinds of fact-checking datasets available and their limitations[6]. The authors have come up with a multi-domain, evidence-based fact-checking dataset. They have also described the metadata of Fact-check articles. [7] describes the task of fact-checking and construction of dataset using the process used by journalists. Several methods are published for automatic detection of fake news, Pérez-Rosas et al. discuss the automatic detection of fake news and linguists differences in false and legitimate content. [9] analyse, compare and summarise the several different methods available for fake news detection. The authors give an overview of methods that have been tested for fake news detection. Research has been done to look inside on the feature of a news story in the modern diaspora along with different kinds of story and their impact on people[10].

Both COVID-19 pandemic and infodemic spread in parallel, Mesquita et al.; proposes a framework to fight against fake medical news because fake news can intensify the effect of the COVID-19 pandemic [11]. All health workers, scientists, the government are trying to fight against fake medical news. In March 2020, several cases were discovered where people are consuming partially true information about COVID-19 on Facebook and WhatsApp without using the proper medical terms, which creates a panic about medicine and prevention for COVID-19 [12]. By analysing the search behaviour of people in Italy from January to March 2020 and found that a "large number of infodemic monikers were observed across Italy" [13]. During the time of the pandemic, fake news is spread all over the world in different languages. The fact-news is covered in different domains like origin and spread, conspiracy theory etc. There is a lack of resources that are multilingual and cross-domain and have been collected from multiple sources.

3. Task Description

The task was organised as CLEF 2021 - CheckThat! Lab Detecting Check-Worthy Claims, Previously Fact-Checked Claims, and Fake News [14, 15]. In task 3 of CheckThat!, the core idea of task is to classify fake news [16]. The task is defined as follows.

3.1 Task Definition

The first task is to classify the text of the news article into the four possible classes into true, partially true, false, or other (e.g., claims in dispute) and predict the topical domain of the article. The task is divided into two sub-task, they are:

Subtask 3A: Multi-class fake news detection of news articles (English)

Sub-task would be the classification of news articles in four classes as defined below:

Subtask 3B: Domain Classification

We often observed that the incoming flow of requests is from different topics such as cancer, COVID-19, diet, etc. With this categorisation, we can assign the respective person/expert to verify the content. It will help to save the manual effort in assigning the claim verification. The task was assigned for the following domain:

4. Experiment

Transfer learning is a technique where a deep learning model trained on a large dataset performs similar tasks on another dataset. We call such a deep learning model a pre-trained model. The most renowned examples of pre-trained models are the computer vision deep learning models trained on the ImageNet dataset. So, it is better to use a pre-trained model as a starting point to solve a problem rather than building a model from scratch.

For training the model, We have to download the external data from the fact-checked articles. We have used the AMUSED framework [17], and FakeCovid [3] concept to merge the articles into four categories. We have downloaded 206,432 facts checked articles from more than 100 fact-checking websites. To gather the external dataset, we developed a Python-based crawler using Beautiful soup [18] to fetch all the news articles from the 92 different fact-checking websites; they are Poynter, Snopes, PolitiFact etc. Our crawler collects important information like the title of the news articles, name of the fact-checking websites, date of publication, the text of the news articles, and a class of news articles. For training, we have used the combination of text and text with the given dataset of the task. We have applied the basic cleaning process of natural language processing like removing special characters, punctuation marks etc.

To classify fake news articles, we used the state-of-the-art neural network language model BERT, which has been pre-trained on a large corpus to solve language processing tasks [19]. An essential advantage of BERT is that it can be fine-tuned for task-specific datasets and allows high text classification accuracy even for smaller datasets. In the context of fake news classification, BERT has already been applied for multiclass classification tasks, for example, on the Chinese social media platform Weibo, where it achieved considerable accuracy[20]. The randomisation of the data prevents seasonal patterns from being learned by the model. The randomisation of the data prevents seasonal patterns from being learned by the model. For BERT fine-tuning, the models for the videos were trained on for four epochs, and the models for the comments were trained on for three epochs with a learning rate of 2e-5. We set hidden units as 300 and training epoch as 200. Each training process continues until the restriction or validation loss is continued. The batch size is set to 10, and the learning rate is 0.001. After the individual prediction on the two test datasets, we evaluated the accuracy of the two models using the weighted F1 score. For classification, the model gave a macro F1 score of 83.76% for Task 3A and 85.55% for Task 3B.

5. Conclusion & Future Work

In this paper, We have presented a classification model for detecting the classification of fake news and its domain. Using the BERT model, the model performed well compared to the traditional machine learning technique. With the proposed model, the problem of fake news classification, domain identification will be addressed. As future work, the concept can be used for determining fake news on a large scale. Even the type of fake news can be embedded in the HTML content of the article [21].

Getting a fake news corpus is a challenging task, and with the limited dataset, it is hard to enhance the machine learning model's performance. With the external dataset, the performance of the classifier is increased. We also observed that different fact-checking websites are checking several duplicates of claims. Several old claims are repurposed as fake news. So, in future work, detecting similar claims would help find fake news classes using text matching.

6. References

[1] L. Graves, F. Cherubini, The rise of fact-checking sites in europe (2016).
[2] G. K. Shahi, A. Dirkson, T. A. Majchrzak, An exploratory study of covid-19 misinformation on twitter, Online social networks and media (2021) 100104.
[3] G. K. Shahi, D. Röchert, S. Stieglitz, Covid ct: Analysis and detection of different conspiracy theories on youtube in the context of covid-19 (2020).
[4] G. K. Shahi, T. A. Majchrzak, Exploring the Spread of COVID-19 Misinformation on Twitter, Technical Report, EasyChair, 2021.
[5] W. Y. Wang, " liar, liar pants on fire": A new benchmark dataset for fake news detection, arXiv preprint arXiv:1705.00648 (2017).
[6] I. Augenstein, C. Lioma, D. Wang, L. C. Lima, C. Hansen, C. Hansen, J. G. Simonsen, Multifc: A real-world multi-domain dataset for evidence-based fact checking of claims, arXiv preprint arXiv:1909.03242 (2019).
[7] A. Vlachos, S. Riedel, Fact checking: Task definition and dataset construction, in: Proceedings of the ACL 2014 Workshop on Language Technologies and Computational Social Science, 2014, pp. 18–22.
[8] V. Pérez-Rosas, B. Kleinberg, A. Lefevre, R. Mihalcea, Automatic detection of fake news, arXiv preprint arXiv:1708.07104 (2017).
[9] X. Zhou, R. Zafarani, Fake news: A survey of research, detection methods, and opportunities, arXiv preprint arXiv:1812.00315 2 (2018).
[10] S. B. Parikh, P. K. Atrey, Media-rich fake news detection: A survey, in: 2018 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR), IEEE, 2018, pp. 436–441.
[11] C. T. Mesquita, A. Oliveira, F. L. Seixas, A. Paes, Infodemia, fake news and medicine: Science and the quest for truth, International Journal of Cardiovascular Sciences (2020).
[12] D. Orso, N. Federici, R. Copetti, L. Vetrugno, T. Bove, Infodemic and the spread of fake news in the covid-19-era, European Journal of Emergency Medicine (2020).
[13] A. Rovetta, A. S. Bhagavathula, Covid-19-related web search behaviors and infodemic attitudes in italy: Infodemiological study, JMIR Public Health and Surveillance 6 (2020) e19374.
[14] P. Nakov, G. Da San Martino, T. Elsayed, A. Barrón-Cedeño, R. Míguez, S. Shaar, F. Alam, F. Haouari, M. Hasanain, N. Babulkov, A. Nikolov, G. K. Shahi, J. M. Struß, T. Mandl, The CLEF-2021 CheckThat! lab on detecting check-worthy claims, previously fact-checked claims, and fake news, in: Proceedings of the 43rd European Conference on Information Retrieval, ECIR '21, Lucca, Italy, 2021, pp. 639–649.
[15] P. Nakov, G. Da San Martino, T. Elsayed, A. Barrón-Cedeño, R. Míguez, S. Shaar, F. Alam, F. Haouari, M. Hasanain, N. Babulkov, A. Nikolov, G. K. Shahi, J. M. Struß, T. Mandl, S. Modha, M. Kutlu, Y. S. Kartal, Overview of the CLEF-2021 CheckThat! lab on detecting check-worthy claims, previously fact-checked claims, and fake news, proceedings of the 12th international conference of the clef association: Information access evaluation meets multiliguality, multimodality, and visualization, series = CLEF '2021, address = Bucharest, Romania (online), 2021.
[16] G. K. Shahi, J. M. Struß, T. Mandl, Overview of the CLEF-2021 CheckThat! lab task 3 on fake news detection, in: Working Notes of CLEF 2021—Conference and Labs of the Evaluation Forum, CLEF '2021, Bucharest, Romania (online), 2021.
[17] G. K. Shahi, Amused: An annotation framework of multi-modal social media data, 2020. arXiv:2010.00502.
[18] L. Richardson, Beautiful soup documentation, April (2007).
[19] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional transformers for language understanding, arXiv preprint arXiv:1810.04805 (2018).
[20] T. Wang, K. Lu, K. P. Chow, Q. Zhu, Covid-19 sensing: negative sentiment analysis on social media in china via bert model, Ieee Access 8 (2020) 138162–138169.
[21] G. K. Shahi, D. Nandini, S. Kumari, Inducing schema.org markup from natural language context, Kalpa Publications in Computing 10 (2019) 38–42.

CLEF 2021 - Conference and Labs of the Evaluation Forum, September 21–24, 2021, Bucharest, Romania

NoFake на CheckThat! 2021: Обнаружение фейковых новостей с использованием BERT

Сушма Кумари
Независимый исследователь
Аннотация

Было проведено много исследований по разоблачению и анализу фейковых новостей. Многие исследователи изучают обнаружение фейковых новостей в последние годы, но многие ограничиваются данными из социальных сетей. В настоящее время множество проверяющих факты публикуют свои результаты в различных форматах. Кроме того, разные проверяющие факты используют различные метки для фейковых новостей, что затрудняет создание универсального классификатора. Объединение классов может улучшить производительность машинной модели. Эта категоризация по доменам поможет группировать статьи, что позволит сэкономить ручные усилия при назначении проверки утверждений. В этой статье мы представили модель классификации на основе BERT для прогнозирования домена и классификации. Мы также использовали дополнительные данные из проверенных статей. Мы достигли макро-F1 оценки 83,76% для Задачи 3A и 85,55% для Задачи 3B с использованием дополнительных обучающих данных.

Ключевые слова: Фейковые новости, Многоклассовая классификация, Классификация доменов, Дезинформация

1. Введение

Данная журналистика - это новая журналистская дисциплина, которая в основном фокусируется на исследованиях, основанных на данных, и форматах представления. Однако фундаментальная проблема дата-журналистики и классической журналистики заключается в том, что многие данные, представляющие журналистский интерес, доступны только в неструктурированной форме: в виде текстов, таблиц и графиков в документах различных типов (Word, PDF, электронная почта и т.д.) или на веб-сайтах. Кроме того, отсутствует централизованный хаб данных для сбора информации из различных источников.

В средствах массовой информации, социальных сетях и других веб-источниках наблюдается растущее количество фейковых новостей. Было проведено много исследований по обнаружению фейковых новостей и их разоблачению [1]. За последние два десятилетия наблюдается огромный рост распространения дезинформации, что также отражается в количестве веб-сайтов по проверке фактов [1]. Веб-сайты по проверке фактов могут помочь в расследовании утверждений и помочь гражданам определить, является ли информация, используемая в статье, правдивой или нет. Более 213 веб-сайтов по проверке фактов работают на 40+ языках в 100+ странах [2, 3, 4].

Не существует стандартного протокола для услуг по проверке фактов среди различных проверяющих, и они не публикуют свои проверенные статьи в стандартном формате, что приводит к нескольким конфликтам. Шахи и др. [3] обсуждают необходимость обнаружения новостных статей, потенциально содержащих ложную информацию. В этой статье мы обсудили метод, использованный для обнаружения фейковых новостей для совместного задания на CheckThat!. Остальная часть статьи организована следующим образом: раздел "Связанные работы" описывает предыдущие работы, "Описание задачи" дает обзор задачи, раздел "Эксперимент" акцентирует внимание на деталях эксперимента, а "Заключение и будущая работа" фокусируется на выводах и будущих аспектах работы.

2. Связанные работы

Проверка фактов - это метод контроля ущерба, который является одновременно необходимым и не масштабируемым. Возможно, будет трудно убрать человеческий компонент из картины в ближайшее время. Но все же автоматическая классификация фейковых новостей может помочь снизить нагрузку на проверяющих факты. Фальшивый контент распространяется в нескольких форматах; многие из них перепрофилируются путем изменения текста, местоположения и т.д.

Обнаружение фейковых новостей - сложная проблема. Исследования проводились по обнаружению фейковых новостей с использованием данных из социальных сетей, таких как твиты [2], видео на YouTube [3], но меньше исследований было сосредоточено на новостных статьях. Одна из основных причин - отсутствие корпуса новостных статей. Фейковые новости распространяются в нескольких областях, таких как преступность, выборы, экономика.

Не хватает корпуса для обучения машинного обучения на основе проверки фактов. За последние несколько лет была опубликована коллекция небольших наборов данных, связанных с проверкой фактов. Эти наборы данных представляют собой смесь тем фейковых новостей, например, выборы в США 2016 года [5]. Тем не менее, проверка фактов доминирует на английском языке и нескольких источниках, таких как Snopes, Politifact и т.д.

Корпус Multi-FC описывает различные виды доступных наборов данных для проверки фактов и их ограничения [6]. Авторы разработали мультидоменный, основанный на доказательствах набор данных для проверки фактов. Они также описали метаданные статей по проверке фактов. [7] описывает задачу проверки фактов и построение набора данных с использованием процесса, применяемого журналистами. Было опубликовано несколько методов для автоматического обнаружения фейковых новостей, Перес-Росас и др. обсуждают автоматическое обнаружение фейковых новостей и лингвистические различия в ложном и легитимном контенте. [9] анализирует, сравнивает и обобщает несколько различных методов, доступных для обнаружения фейковых новостей. Авторы дают обзор методов, которые были протестированы для обнаружения фейковых новостей. Были проведены исследования, чтобы заглянуть внутрь особенностей новостных историй в современной диаспоре, а также различных видов историй и их влияния на людей [10].

Пандемия COVID-19 и инфодемия распространялись параллельно, Мескита и др. предлагают структуру для борьбы с фальшивыми медицинскими новостями, потому что фейковые новости могут усиливать эффект пандемии COVID-19 [11]. Все медицинские работники, ученые, правительство пытаются бороться с фальшивыми медицинскими новостями. В марте 2020 года было обнаружено несколько случаев, когда люди потребляли частично правдивую информацию о COVID-19 в Facebook и WhatsApp без использования надлежащих медицинских терминов, что создавало панику по поводу лекарств и профилактики COVID-19 [12]. Анализируя поисковое поведение людей в Италии с января по март 2020 года, было обнаружено, что "большое количество инфодемических прозвищ наблюдалось по всей Италии" [13]. Во время пандемии фейковые новости распространяются по всему миру на разных языках. Факт-новости охватываются в различных областях, таких как происхождение и распространение, теории заговора и т.д. Не хватает ресурсов, которые являются многоязычными и кроссплатформенными и были собраны из нескольких источников.

3. Описание задачи

Задача была организована как CLEF 2021 - CheckThat! Lab Detecting Check-Worthy Claims, Previously Fact-Checked Claims, and Fake News [14, 15]. В задаче 3 CheckThat! основная идея задачи - классифицировать фейковые новости [16]. Задача определяется следующим образом.

3.1 Определение задачи

Первая задача - классифицировать текст новостной статьи на четыре возможных класса: правда, частично правда, ложь или другое (например, спорные утверждения) и предсказать тематическую область статьи. Задача разделена на две подзадачи:

Подзадача 3A: Многоклассовое обнаружение фейковых новостей в новостных статьях (английский язык)

Подзадача заключается в классификации новостных статей на четыре класса, определенных ниже:

Подзадача 3B: Классификация доменов

Мы часто наблюдали, что входящий поток запросов поступает из разных тем, таких как рак, COVID-19, диета и т.д. С такой категоризацией мы можем назначить соответствующего человека/эксперта для проверки содержания. Это поможет сэкономить ручные усилия при назначении проверки утверждений. Задача была назначена для следующих доменов:

4. Эксперимент

Перенос обучения - это техника, при которой модель глубокого обучения, обученная на большом наборе данных, выполняет аналогичные задачи на другом наборе данных. Мы называем такую модель глубокого обучения предварительно обученной моделью. Наиболее известными примерами предварительно обученных моделей являются модели глубокого обучения компьютерного зрения, обученные на наборе данных ImageNet. Поэтому лучше использовать предварительно обученную модель в качестве отправной точки для решения проблемы, чем строить модель с нуля.

Для обучения модели нам необходимо загрузить внешние данные из проверенных статей. Мы использовали структуру AMUSED [17] и концепцию FakeCovid [3] для объединения статей в четыре категории. Мы загрузили 206 432 проверенных фактами статей с более чем 100 веб-сайтов по проверке фактов. Для сбора внешнего набора данных мы разработали краулер на Python с использованием Beautiful soup [18] для получения всех новостных статей с 92 различных веб-сайтов по проверке фактов; это Poynter, Snopes, PolitiFact и т.д. Наш краулер собирает важную информацию, такую как заголовок новостных статей, название веб-сайтов по проверке фактов, дата публикации, текст новостных статей и класс новостных статей. Для обучения мы использовали комбинацию текста и текста с предоставленным набором данных задачи. Мы применили базовый процесс очистки естественного языка, такой как удаление специальных символов, знаков пунктуации и т.д.

Для классификации фейковых новостных статей мы использовали современную языковую модель нейронной сети BERT, которая была предварительно обучена на большом корпусе для решения задач обработки языка [19]. Важным преимуществом BERT является то, что его можно дообучить для наборов данных, специфичных для задачи, и он позволяет достичь высокой точности классификации текста даже для небольших наборов данных. В контексте классификации фейковых новостей BERT уже применялся для задач многоклассовой классификации, например, на китайской платформе социальных сетей Weibo, где он достиг значительной точности [20]. Рандомизация данных предотвращает изучение сезонных закономерностей моделью. Для тонкой настройки BERT модели для видео обучались в течение четырех эпох, а модели для комментариев обучались в течение трех эпох со скоростью обучения 2e-5. Мы установили скрытые единицы равными 300, а эпоху обучения - 200. Каждый процесс обучения продолжается до тех пор, пока не будет достигнуто ограничение или потеря на валидации не продолжится. Размер пакета установлен на 10, а скорость обучения - 0,001. После индивидуального прогнозирования на двух тестовых наборах данных мы оценили точность двух моделей с использованием взвешенной оценки F1. Для классификации модель дала макро-F1 оценку 83,76% для Задачи 3A и 85,55% для Задачи 3B.

5. Заключение и будущая работа

В этой статье мы представили модель классификации для обнаружения классификации фейковых новостей и их домена. С использованием модели BERT модель показала хорошие результаты по сравнению с традиционной техникой машинного обучения. С предложенной моделью будет решена проблема классификации фейковых новостей, идентификации домена. В качестве будущей работы концепция может быть использована для определения фейковых новостей в больших масштабах. Даже тип фейковых новостей может быть встроен в HTML-контент статьи [21].

Получение корпуса фейковых новостей - сложная задача, и с ограниченным набором данных трудно улучшить производительность модели машинного обучения. С внешним набором данных производительность классификатора увеличивается. Мы также заметили, что разные веб-сайты по проверке фактов проверяют несколько дубликатов утверждений. Несколько старых утверждений перепрофилируются как фейковые новости. Таким образом, в будущей работе обнаружение похожих утверждений поможет найти классы фейковых новостей с использованием сопоставления текста.

6. Список литературы

[1] L. Graves, F. Cherubini, The rise of fact-checking sites in europe (2016).
[2] G. K. Shahi, A. Dirkson, T. A. Majchrzak, An exploratory study of covid-19 misinformation on twitter, Online social networks and media (2021) 100104.
[3] G. K. Shahi, D. Röchert, S. Stieglitz, Covid ct: Analysis and detection of different conspiracy theories on youtube in the context of covid-19 (2020).
[4] G. K. Shahi, T. A. Majchrzak, Exploring the Spread of COVID-19 Misinformation on Twitter, Technical Report, EasyChair, 2021.
[5] W. Y. Wang, " liar, liar pants on fire": A new benchmark dataset for fake news detection, arXiv preprint arXiv:1705.00648 (2017).
[6] I. Augenstein, C. Lioma, D. Wang, L. C. Lima, C. Hansen, C. Hansen, J. G. Simonsen, Multifc: A real-world multi-domain dataset for evidence-based fact checking of claims, arXiv preprint arXiv:1909.03242 (2019).
[7] A. Vlachos, S. Riedel, Fact checking: Task definition and dataset construction, in: Proceedings of the ACL 2014 Workshop on Language Technologies and Computational Social Science, 2014, pp. 18–22.
[8] V. Pérez-Rosas, B. Kleinberg, A. Lefevre, R. Mihalcea, Automatic detection of fake news, arXiv preprint arXiv:1708.07104 (2017).
[9] X. Zhou, R. Zafarani, Fake news: A survey of research, detection methods, and opportunities, arXiv preprint arXiv:1812.00315 2 (2018).
[10] S. B. Parikh, P. K. Atrey, Media-rich fake news detection: A survey, in: 2018 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR), IEEE, 2018, pp. 436–441.
[11] C. T. Mesquita, A. Oliveira, F. L. Seixas, A. Paes, Infodemia, fake news and medicine: Science and the quest for truth, International Journal of Cardiovascular Sciences (2020).
[12] D. Orso, N. Federici, R. Copetti, L. Vetrugno, T. Bove, Infodemic and the spread of fake news in the covid-19-era, European Journal of Emergency Medicine (2020).
[13] A. Rovetta, A. S. Bhagavathula, Covid-19-related web search behaviors and infodemic attitudes in italy: Infodemiological study, JMIR Public Health and Surveillance 6 (2020) e19374.
[14] P. Nakov, G. Da San Martino, T. Elsayed, A. Barrón-Cedeño, R. Míguez, S. Shaar, F. Alam, F. Haouari, M. Hasanain, N. Babulkov, A. Nikolov, G. K. Shahi, J. M. Struß, T. Mandl, The CLEF-2021 CheckThat! lab on detecting check-worthy claims, previously fact-checked claims, and fake news, in: Proceedings of the 43rd European Conference on Information Retrieval, ECIR '21, Lucca, Italy, 2021, pp. 639–649.
[15] P. Nakov, G. Da San Martino, T. Elsayed, A. Barrón-Cedeño, R. Míguez, S. Shaar, F. Alam, F. Haouari, M. Hasanain, N. Babulkov, A. Nikolov, G. K. Shahi, J. M. Struß, T. Mandl, S. Modha, M. Kutlu, Y. S. Kartal, Overview of the CLEF-2021 CheckThat! lab on detecting check-worthy claims, previously fact-checked claims, and fake news, proceedings of the 12th international conference of the clef association: Information access evaluation meets multiliguality, multimodality, and visualization, series = CLEF '2021, address = Bucharest, Romania (online), 2021.
[16] G. K. Shahi, J. M. Struß, T. Mandl, Overview of the CLEF-2021 CheckThat! lab task 3 on fake news detection, in: Working Notes of CLEF 2021—Conference and Labs of the Evaluation Forum, CLEF '2021, Bucharest, Romania (online), 2021.
[17] G. K. Shahi, Amused: An annotation framework of multi-modal social media data, 2020. arXiv:2010.00502.
[18] L. Richardson, Beautiful soup documentation, April (2007).
[19] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional transformers for language understanding, arXiv preprint arXiv:1810.04805 (2018).
[20] T. Wang, K. Lu, K. P. Chow, Q. Zhu, Covid-19 sensing: negative sentiment analysis on social media in china via bert model, Ieee Access 8 (2020) 138162–138169.
[21] G. K. Shahi, D. Nandini, S. Kumari, Inducing schema.org markup from natural language context, Kalpa Publications in Computing 10 (2019) 38–42.

CLEF 2021 - Конференция и лаборатории форума оценки, 21–24 сентября 2021 года, Бухарест, Румыния