The Fact-Checking Observatory – Reporting the Co-Spread of Misinformation and Fact-checks on Social Media
Authors: Grégoire Burel, Harith Alani
Published in HT '23: 34th ACM Conference on Hypertext and Social Media, Rome, Italy, September 4-8, 2023 · DOI: 10.1145/3603163.3609042 · License: © Copyright held by the owner/author(s).
Abstract
In the context of the Covid-19 pandemic and the Russian invasion of Ukraine, tracking how misinformation and fact-checks spread on social media is key for understanding where fact-checking efforts need to be focused and what demographics are most likely to spread misinformation. In this article, we introduce the Fact-checking Observatory, a website that automatically generates human-readable weekly reports about the spread of misinformation and fact-checks on Twitter. The proposed approach differs from other tools that give one-off manual reports or visualisation by providing organisations and individuals with easily readable and shareable self-contained reports that contain both information about the spread of misinformation and fact-checks.
CCS CONCEPTS
• Information systems →Web mining; Social networking sites.
Fact-checking, Misinformation, Social media, Monitoring, Automatic Reporting.
1 INTRODUCTION
Recent research indicates that misinformation on social media tends to spread much faster than true information [5]. In this context, access to corrective information such as the information provided by fact-checking organisations1 is key for fighting the spread of misinformation. The Fact-checking Observatory2 (FCO) is a website that aims at helping fact-checking organisations and the public to learn how effective fact-checking articles are, and supporting the identification of important topics and demographics susceptible to spreading
1International Fact-Checking Network (IFCN), https://www.poynter.org/ifcn. 2Fact-checking Observatory, https://fcobservatory.org.
misinformation. The FCO complements existing research on the understanding of misinformation and fact-checks co-spread during the Covid-19 pandemic[1, 2] by providing a service that generates weekly reports about misinformation risk areas and indicators about the effectiveness of fact-checking on Twitter3 in the context of the Covid-19 pandemic and the Russian invasion of Ukraine. The main impact of the FCO is the creation of easily readable and up-to-date reports based on previous research on misinformation spread and fact-checking effectiveness [1, 2]. The reports are designed to be actionable so that fact-checking organisations as well as journalists and authorities can use them to understand how useful fact-checking articles are for fighting misinformation spread during Covid-19 and the Russian invasion of Ukraine. Rather than providing a single contentiously updated report, the FCO opts for the generation of weekly reports. We argue that this approach is complementary to other methods as it allows organisations and individuals to better follow the evolution of misinformation and fact-checks spread while providing shareable self-contained reports. Another key difference of the FCO against existing misinformation observatory is that it both reports on the spread of misinformation and fact-checks.
2 WEEKLY AUTOMATIC CO-SPREAD REPORTING
The goal of the FCO is to create human-readable reports that are not created using any manual input. For generating the reports, we use templates that are filled every week as new data is collected. For transparency purposes, templates and data collection is versioned so readers are aware when new data and new templates are used for a given report. In order to not have reports always look exactly the same, some parts of the reports are generated using different variations of the same sentence at random. A report example is displayed in Figure 1. The FCO relies on a fact-checking database collected from organisations that are vetted by the IFCN. The database contains misinformation and fact-check URL pairs that are then searched for on Twitter. The collected tweets are fed to machine learning models to extract user demographics [6] (i.e., gender, organisation, language, and age group). The tweets and demographic information are then used to generate the FCO weekly reports. The database for the Covid-related reports comes from the CoronaVirusFacts/ DatosCoronaVirus Alliance Database4 whereas the Russian invasion database is obtained from the MisinfoMe tool database [4]. Currently (22/05/2023), the FCO has published 156 reports about Covid-19 and 79 reports about the Russian invasion of Ukraine. The
3Twitter, https://twitter.com. 4IFCN coronavirus database, https://www.poynter.org/ifcn-covid-19-misinformation.
FCO reports are divided into 5 different sections summarising the evolution of misinformation compared to the previous report and the state of misinformation and fact-checking spread:
(1) Title and subtitle: The first part of a report contains the title and subtitle that is automatically generated at random based on what topic spread increased or decreased the most during the reporting period compared to the previous report. This section also includes the version of the report and data used for generating the web page so that readers are aware if a report or underlying data has changed.
(2) Summary statistics and disclaimer: The second section contains statistics concerning the data used for the report as well as information about how the report is generated. The summary section gives details about the amount of misinforming tweets collected so far, the number of additional misinformation collected since the previous report and the difference between the current report spread increase and the previous report spread. Similar statistics are also reported for the amount of shared fact-checking URLs as well as the number of fact-check URLs used for generating the data so far and the number of fact-checking organisations involved. This section gives a glance at how misinformation or fact-checks spread differs from the previous reporting period.
(3) Key content, topics and provenance: For the Covid-19 reports, the number of posts shared over time about covid-related topics and statistics about the topics that were the most and least shared during the reporting period are given for quickly identifying key topics and posts during the reporting period. For the Russian invasion reports, the sources of the URLs tracked by the FCO are given since topic information is not available.
(4) Fact-checking and spreaders location: For the Covid-19 reports, information about who is providing the fact-checks such as the language of the fact-checking organisation and their location and statistics about what topic spreads the most for both misinformation and fact-checks is also reported. For the Russian invasion reports, the estimated location of the user spreading fact-checks and misinformation is reported.
(5) Locations and mentions (Russian invasion reports only): For the Russian invasion report, we also extract [3] key locations and individuals mentioned in the fact-checking articles and report on their importance.
(6) Demographic impact: The last section is based on the automatic extraction of demographic information [6] about who shares misinformation and fact-checks with details about which gender, age group and account type spreads misinformation and fact-checks the most. This can be used for better directing fact-checking efforts towards particular demographics or identifying what demographic is the most sensible to misinformation.
Acknowledgments
This work has received support from the European Union's Horizon 2020 research and innovation programme under grants agreement No 101003606 (HERoS).
Figure 1: A Covid-19 Fact-Checking Observatory report example for the 11-18 April 2022 period.
References
[1] Grégoire Burel, Tracie Farrell, and Harith Alani. 2021. Demographics and topics impact on the co-spread of COVID-19 misinformation and fact-checks on Twitter.
Information Processing & Management 58, 6 (November 2021). https://oro.open.ac. uk/78748/
[2] Grégoire Burel, Tracie Farrell, Martino Mensio, Prashant Khare, and Harith Alani. 2020. Co-spread of Misinformation and Fact-Checking Content During the Covid19 Pandemic. In Social Informatics, Samin Aref, Kalina Bontcheva, Marco Braghieri, Frank Dignum, Fosca Giannotti, Francesco Grisolia, and Dino Pedreschi (Eds.). Springer International Publishing, Cham, 28–42.
[3] Antonin Delpeuch. 2019. Opentapioca: Lightweight entity linking for wikidata. arXiv preprint arXiv:1904.09131 (2019).
[4] Martino Mensio, Grégoire Burel, Tracie Farrell, and Harith Alani. 2023. MisinfoMe: A Tool for Longitudinal Assessment of Twitter Account's Sharing of Misinformation. In Proceedings of the 2023 ACM Conference on User Modeling, Adaptation and
Personalization (Singapore, Singapore) (UMAP '23). Association for Computing Machinery, Limassol, Cyprus.
[5] Soroush Vosoughi, Deb Roy, and Sinan Aral. 2018. The spread of true and false news online. Science 359, 6380 (2018), 1146–1151.
[6] Zijian Wang, Scott Hale, David Ifeoluwa Adelani, Przemyslaw Grabowicz, Timo Hartman, Fabian Flöck, and David Jurgens. 2019. Demographic Inference and Representative Population Estimates from Multilingual Social Media Data. In The World Wide Web Conference (San Francisco, CA, USA) (WWW '19). Association for Computing Machinery, New York, NY, USA, 2056–2067. https://doi.org/10. 1145/3308558.3313684
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime