EarlyAd: A System for Real-Time Surveillance of Brazilian Early Electoral Ads on Twitter
Marcelo M. R. Araújo Samuel Guimarães Márcio Silva Universidade Federal de Minas Gerais Universidade Federal de Minas Gerais Universidade Federal de Minas Gerais, Belo Horizonte, Minas Gerais, Brazil Belo Horizonte, Minas Gerais, Brazil Universidade Federal de Mato Grosso marceloaraujo@dcc.ufmg.br samuelsg@ufmg.br do Sul Belo Horizonte, Minas Gerais, Brazil marcio@facom.ufms.br Josemar Caetano Jonatas Santos Julio C. S. Reis Universidade Federal de Minas Gerais Universidade Federal de Minas Gerais Universidade Federal de Viçosa Belo Horizonte, Minas Gerais, Brazil Belo Horizonte, Minas Gerais, Brazil Belo Horizonte, Minas Gerais, Brazil josemarcaetano@dcc.ufmg.br jonatashds@ufmg.br jreis@ufv.br Ana P. C. Silva Fabrício Benevenuto Jussara M. Almeida Universidade Federal de Minas Gerais Universidade Federal de Minas Gerais Universidade Federal de Minas Gerais Belo Horizonte, Minas Gerais, Brazil Belo Horizonte, Minas Gerais, Brazil Belo Horizonte, Minas Gerais, Brazil ana.coutosilva@dcc.ufmg.br fabricio@dcc.ufmg.br jussara@dcc.ufmg.br
ABSTRACT
The sheer volume of social media data produced daily brings challenges to the surveillance and detection of specific actions of interest, notably actions associated with infringement of regulations or laws in force in specific countries. One such case is the sharing of early electoral advertisement (ads) , which is prohibited by law during political elections in Brazil. In this paper, we introduce EarlyAd, a system that performs, in real time, the collection, identification and analysis of early electoral advertisements on Twitter. Our tool was designed to bring more transparency and awareness to the Brazilian society about such practice, uncovering common patterns associated with this kind of content, while also offering competent authorities evidence to facilitate law enforcement. Our tool is running online and it is available at: http://earlyad.dcc.ufmg.br.
CCS CONCEPTS
• Human-centered computing → Collaborative and social computing systems and tools .
KEYWORDS
Twitter, Politics, Early Electoral Ads, Political Advertisement, Propaganda
ACM Reference Format:
Marcelo M. R. Araújo, Samuel Guimarães, Márcio Silva, Josemar Caetano, Jonatas Santos, Julio C. S. Reis, Ana P. C. Silva, Fabrício Benevenuto, and Jussara M. Almeida. 2022. EarlyAd: A System for Real-Time Surveillance of Brazilian Early Electoral Ads on Twitter. In Proceedings of the 33rd ACM
Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). HT ’22, June 28-July 1, 2022, Barcelona, Spain
© 2022 Copyright held by the owner/author(s). ACM ISBN 978-1-4503-9233-4/22/06. https://doi.org/10.1145/3511095.3536376
Conference on Hypertext and Social Media (HT ’22), June 28-July 1, 2022, Barcelona, Spain. ACM, New York, NY, USA, 3 pages. https://doi.org/10. 1145/3511095.3536376
1 INTRODUCTION
The ever-increasing use of social media platforms as prime source of information has raised concerns regarding the monitoring and identification of sensitive content sharing such as hate speech, disinformation and, more generally, any content whose sharing infringes laws and regulations in force [6]. In this work, we focus on a particular type of (possibly sensitive) content, political advertisement (ads) , whose sharing is constrained by law in some countries such as Brazil.
In Brazil, it is prohibited by law to declare candidacy in a political election and to make any (explicit or implicit) request to vote ahead of time. That is, electoral advertisement can only take place during a predefined period preceding election date. Such constraint is aimed at guaranteeing the equality of treatment among candidates, a basic principle of the electoral process in the country. The Superior Electoral Court (TSE) is the governmental institution responsible for establishing such period and enforcing it is respected by all individuals (candidates or not). The sharing of electoral ads in any media platform (including social media) before the period defined by TSE, referred to as early electoral ads , is a crime provided for by Law nº 9.504/1997 in the Brazilian Electoral Code [14]. Traditional media (e.g., broadcast television or radio) has extensive supervision against early electoral ads, but the same is not true for social media platforms such as Twitter. Indeed, these environments have become fertile ground for political debate and, inevitably for the spread of such content [2, 9], which has raised concerns from TSE and society in general.
In that context, we have used a collection of different technologies to build a real-time monitor, called EarlyAd , to detect the sharing of early electoral ads on Twitter. The main goal of our tool is to provide transparency to the Brazilian society about such practice,
HT ’22, June 28-July 1, 2022, Barcelona, Spain
Marcelo Araújo, et al.
uncovering common patterns associated with this kind of content. By doing so, EarlyAd also contributes to raise awareness in the population, while also offering to competent authorities potential evidence of criminal actions by exposing content (i.e., tweets) spreading electoral ads before the allowed period. We do not disclose any information about the tweet’s authors, only the tweet ID, following Twitter Private Information and Media Policy [15].
Our EarlyAd system has three main components, namely: (1) Data collection, which comprises a Twitter collector to gather tweets associated with pre-established keywords in real-time (i.e., working in streaming mode), (2) a supervised classifier to identify tweets containing electoral ads associated with ongoing political election campaigns, and (3) a web interface enabling various analyses of the classified tweets and offering visualizations of the data to the end-user. The tool is running online and, in the interest of contributing to the transparency of the electoral process, is accessible to everyone at http://earlyad.dcc.ufmg.br. In the next section, we discuss its architecture.
2 SYSTEM ARCHITECTURE
2.1 Data Collection
We designed a crawler to continuously gather tweets using the streaming Twitter API<sup>1</sup> . Our crawler was implemented in Python, using the library TwitterAPI<sup>2</sup> , and requires an API key (which must be provided as input to the system in the web interface) to be executed<sup>3</sup> . Our crawler collects text from tweets based on a dictionary of keywords (in Portuguese) related to electoral ads<sup>4</sup> . This dictionary was originally suggested by a domain expert and further refined to include other terms based on early observations. Examples of keywords are (in English): ’Vote for president’, ’I want to be your candidate’ and ’if I am elected’ . Since our crawler gathers tweets in real time, towards detecting early electoral ads, it should be executed before the electoral periods defined by the TSE starts. Otherwise, if run during the allowed period it could still be useful, for transparency purposes, by offering visualizations about the (legal) electoral ads being shared on Twitter.
2.2 Classification of Electoral Ads
Prior work on analyzing political content in social media relies either on supervised machine learning (ML) models [5] or unsupervised lexical approaches [12]. Most prior efforts focus on political content in general, while only a few aimed at identifying electoral ads [1, 8, 10, 11] but these focused on other platforms (Facebook) or content in English. The identification of electoral ads in Portuguese on Twitter is challenging due to: (1) restrictions of tools for the language (i.e., limited availability of natural language processing tools) [3]; (2) differences in narratives across platforms (e.g., more concise content on Twitter) and (3) large presence of tweets related to other (non-political) voting processes that follow similar textual patterns. Aiming at designing a classification model to identify tweets containing electoral ads, we experimented with a dictionary-based and
1https://developer.twitter.com/en/docs/twitter-api
2https://github.com/geduldig/TwitterAPI
3https://developer.twitter.com/en/docs/twitter-api/getting-started/getting-access-tothe-twitter-api
4In this first version of our system, we consider only textual content, thus images and videos associated with tweets are discarded.
various machine learning (ML) methods. The former explored a previously built dictionary of keywords extended with a blacklist of terms to distinguish non-political election-related content. Regarding ML methods, we experimented with traditional approaches (SVM, Naive Bayes, and Logistic Regression), tree-based classifiers (Random Forest and Gradient Boosting) as well as deep learning methods (CNN, LSTM, RNN, and HAN) [4]. In all cases, given an input tweet, the classifier outputs a score, interpreted as the chance that the tweets contain electoral ads. Given the great degree of subjectivity associated with identifying (especially implicit) electoral ads [13], we believe that an accurate detection will demand human assessment. These scores are meant to assist competent authorities to better direct their (detection) efforts to the more suspicious tweets.
To train the supervised ML models, we collected datasets containing tweets from three early electoral periods from past elections in Brazil, namely from January to August of 2016, 2018 and from January to September 2020. The collection was driven by the aforementioned dictionary of keywords. All tweets gathered were manually labeled by three volunteers into the ones containing electoral ads or those which did not. The agreement among volunteers was moderate (Fleiss’ Kappa κ =0.53 [7]), which is acceptable given the large presence of subjective and implicit ads. Using the labeled datasets, we evaluated the considered nine ML methods in different scenarios, varying the source of training and test sets. For instance, we considered both training and test sets composed of tweets from the selected electoral period as well as from different periods (training coming from the early period). Based on our experiments, we found that the RNN model<sup>5</sup> as set to drive periodic retraining efforts regarding the classification methods.
2.3 Web Interface
The analyses provided by our EarlyAd system are displayed to the final users through a friendly web interface as presented in the narrated screen capture available at https://bit.ly/3FF73rn. Overall, our analyses are built on two complementary views: (i) coarsegrained, which presents a dataset overview and; (ii) fine-grained, which permits one to look individually into the collected tweets. Next, we describe each of them.
In the coarse-grained view, EarlyAd allows one to explore basic information extracted from the dataset, aiming to provide details about what has already been collected and some analysis generated with all the stored tweets. It displays some information such as the total amount of tweets collected as well as the number of tweets collected in the last day and in the last week.
The tool also provides the total number of unique authors in the dataset, and the top-15 users who shared tweets with the chance of carrying electoral ads above a given threshold. Through the analysis of their tweets and the communication patterns of those users, one may have some some insights into how this type of information spreads on Twitter. Finally, one can track how electoral ads score changes daily, weekly, monthly, and yearly, helping users to identify in which periods there was a higher incidence of electoral
5Sequential model using word2vec word embedding (300 dim.), followed by four GRU hidden layers, and dense output layer with softmax function. Training over 10 epochs and 10- fold cross-validation.
HT ’22, June 28-July 1, 2022, Barcelona, Spain
EarlyAd: A System for Real-Time Surveillance of Brazilian Early Electoral Ads on Twitter
advertisements, directing them to look for external events or news that could explain these variations.
Fine-grained view is based on the use of filters to allow one zooming in the collected tweets. The main filters are: the text search, score, and date. These filters permit the user to query and analyze specific subset of tweets that they might want to investigate. EarlyAd also provides word clouds and co-occurrence graphs, helping the user to visualize the tweets’ top-k terms and how they are correlated without having to analyze each tweet individually.
By default, tweets with the highest electoral scores are displayed to the user. However, it is also possible to navigate through the tweets with lower scores, allowing users to visualize all tweets in the database, since the classification method does not assure the presence or absence of electoral advertisements in a tweet. In the future, we plan to add a user feedback mechanism to analyse how different opinions and insights into what constitutes an early political ad may improve the accuracy of the classification task. In other words, we plan to introduce the human-in-the-loop paradigm in EarlyAd. Finally, if one wants to further explore the characteristics of a particular tweet, the tool redirects the user to the tweet page on Twitter, providing information about the tweet’s author, the tweet content, the context of the conversation, and tweets with similar content as well.
Our system was developed as a web-based application. The frontend component was developed using HTML, CSS and JavaScript, while the back-end was developed using Python (version 3.9.7) and the Django Framework to integrate the different parts of the stack development. We note that in respect to user privacy, we do not share or disclose any Personally Identifiable Information (PII) such as Twitter identifier, @username, and other data that could directly identify a particular user. Last, the code to the repository can be found in the following link: https://bit.ly/3N9GRrf.
3 EXECUTION REQUIREMENTS
To host EarlyAd, it is recommended a Linux server with at least 32 GB of RAM, once we used a deep-learning model for the classification task, which demands much memory. Furthermore, it is also required to have the Twitter API key, to enable the data collector (it must be provided as input in the web interface).
REFERENCES
[1] Athanasios Andreou, Giridhari Venkatadri, Oana Goga, Krishna P Gummadi, Patrick Loiseau, and Alan Mislove. 2018. Investigating ad transparency mechanisms in social media: A case study of Facebook’s explanations. In The Network and Distributed System Security Symposium (NDSS) .
[2] Michael Conover, Jacob Ratkiewicz, Matthew Francisco, Bruno Gonçalves, Filippo Menczer, and Alessandro Flammini. 2011. Political polarization on twitter. In Proc. of the ICWSM .
[3] João Ferreira, Hugo Gonçalo Oliveira, and Ricardo Rodrigues. 2019. Improving NLTK for Processing Portuguese. In 8th Symposium on Languages, Applications and Technologies (SLATE 2019) (OpenAccess Series in Informatics (OASIcs), Vol. 74) , Ricardo Rodrigues, Jan Janousek, Luís Ferreira, Luísa Coheur, Fernando Batista, and Hugo Gonçalo Oliveira (Eds.). Schloss Dagstuhl–Leibniz-Zentrum fuer Informatik, Dagstuhl, Germany, 18:1–18:9. https://doi.org/10.4230/OASIcs.SLATE. 2019.18
[4] Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016. Deep Learning . MIT Press. http://www.deeplearningbook.org.
[5] Justin Grimmer and Brandon M. Stewart. 2013. Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts. Political Analysis 21, 3 (2013), 267–297. https://doi.org/10.1093/pan/mps028
[6] Jennifer Jerit and Yangzi Zhao. 2020. Political Misinformation. Annual Review of Political Science 23, 1 (2020), 77–94. https://doi.org/10.1146/annurev-polisci050718-032814
[7] J R Landis and G G Koch. 1977. The measurement of observer agreement for categorical data. Biometrics 33, 1 (March 1977), 159–174.
[8] Victor Le Pochat, Laura Edelson, Tom Van Goethem, Wouter Joosen, Damon McCoy, and Tobias Lauinger. 2022. An Audit of Facebook’s Political Ad Policy Enforcement. In 31st USENIX Security Symposium (USENIX Security 22) . USENIX Association, Boston, MA. https://www.usenix.org/conference/usenixsecurity22/ presentation/lepochat
[9] Filipe N. Ribeiro, Koustuv Saha, Mahmoudreza Babaei, Lucas Henrique, Johnnatan Messias, Fabricio Benevenuto, Oana Goga, Krishna P. Gummadi, and Elissa M. Redmiles. 2019. On Microtargeting Socially Divisive Ads: A Case Study of RussiaLinked Ad Campaigns on Facebook. In Proc. of the FAT .
[10] Márcio Silva and Fabrício Benevenuto. 2021. COVID-19 Ads as Political Weapon. In 36th ACM/SIGAPP Symposium On Applied Computing (SAC’21).
[11] Márcio Silva, Lucas Santos de Oliveira, Athanasios Andreou, Pedro Olmo Vaz de Melo, Oana Goga, and Fabrício Benevenuto. 2020. Facebook Ads Monitor: An Independent Auditing System for Political Ads on Facebook. In Proceedings of The Web Conference 2020 . 224–234.
[12] Tamara A Small. 2011. What the hashtag? A content analysis of Canadian politics on Twitter. Information, Communication & Society 14, 6 (2011), 872–895.
[13] Vera Sosnovik and Oana Goga. 2021. Understanding the complexity of detecting political ads. In Proceedings of the Web Conference 2021 . 2002–2013.
[14] Superior Electoral Court (TSE). 1997. Electoral Law . https://www.tse.jus.br/ legislacao/codigo-eleitoral/lei-das-eleicoes/lei-das-eleicoes-lei-nb0-9.504-de30-de-setembro-de-1997
[15] Twitter. 2022. Private information and media policy. https://help.twitter.com/en/rules-and-policies/personal-information. Accessed: 2022-04-15.
4 CONCLUDING REMARKS
The control of early electoral ads aims to preserve the isonomy between the candidates in the Brazilian electoral process, avoiding some candidates (or their supporters) starting to publicize their campaign before their pairs. To help with this task, we proposed EarlyAd, which is a real-time tool to monitor this phenomenon on Twitter. Our tool allows the investigation of the patterns associated with early electoral ads shared on this platform. We believe that the use of our tool provides transparency to Brazilian society about this phenomenon, allowing any interested party to investigate it.
ACKNOWLEDGMENTS:
This work was partially supported by research grants from Ministério Público de Minas Gerais, CNPq, FAPEMIG, and FAPESP.
---
Source and attribution
Complete full-text transcription and rendered visual material from the ACM Hypertext and Social Media 2022 proceedings PDF. Original publication: ACM, Proceedings of the 33rd ACM Conference on Hypertext and Social Media (HT ’22), Barcelona, Spain, June 28–July 1, 2022. DOI: 10.1145/3511095.3536376.
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime