Abstract
War discourse on YouTube emerges through linked interactions among videos, channels, audiences, comments, and replies. This paper compares YouTube discussions of two U.S.-linked conflicts: the Afghanistan withdrawal in 2021 and the Iran conflict escalation in 2026. We analyze four corpora separating news organizations from political influencers, comprising 340,383 comments from 212,968 unique commenters. Our framework combines channel-level discourse embeddings adapted from WEAT-style semantic association tests, and classifier-based toxicity and hate-related proxy scores. The results show that Afghanistan and Iran discussions form distinct discourse spaces, while some ideologically different actors converge toward similar patterns of conflict commentary. Semantic associations around geopolitical terms vary across conflicts and source types, indicating that the same vocabulary is framed differently across communities. Hostility analysis further shows that influencer-centered spaces tend to attract more toxic discussion, while the Iran corpus exhibits stronger and more persistent hate-related signals, including in replies.
1 Introduction
Online discussions of war emerge through interconnected relations among media producers, videos, audiences, comments, and replies. These structures shape not only which information receives attention, but also how geopolitical actors and events are interpreted, evaluated, and contested. This is particularly relevant in contemporary conflicts, as traditional news organizations now share the information environment with political influencers who often provide more personal, emotional, or explicitly ideological interpretations. In periods of heightened polarization, these networked discussions may also intensify hostile narratives and hateful expression [19, 23, 24].
In this paper, we compare YouTube discourse surrounding two U.S.-linked conflicts: the final stage of the war in Afghanistan in 2021 and the escalation of direct conflict between the United States and Iran in 2026. The cases represent distinct geopolitical moments separated by five years of change in online political communication. We analyze news organizations and political influencers separately because they constitute different forms of information intermediation and attract different audience communities. YouTube is especially suitable for this comparison because it preserves visible links among channels, videos, top-level comments, and replies, allowing conflict narratives and audience reactions to be examined within the same conversational environment.
We investigate three questions: (RQ1) how channel-centered discourse communities differ across conflicts and source types; (RQ2) how these communities semantically associate geopolitical actors and conflict-related concepts; and (RQ3) how hostile and hate-related discourse varies across news and influencer spaces and between top-level comments and replies. We construct four corpora—Afghanistan–News, Afghanistan–Influencers, Iran–News, and Iran–Influencers—comprising 340,383 comments from 212,968 unique commenters. To address RQ1, we adapt POLAR [3] to learn representations of the comment communities surrounding each channel in each conflict. To address RQ2, we use a shared embedding space and WEAT-style association scores [5] to compare the evaluative orientation of geopolitical terms across corpora. Finally, for RQ3, we use continuous toxicity and hate-related classifier scores as comparative proxies for antagonistic discourse.
Our results show that conflict context is the strongest organizing dimension of the channel-community embedding space: Afghanistan and Iran form clearly distinct regions, whereas news and influencer channels show little global separation by source type. Nevertheless, some politically different actors converge across conflicts, indicating that discourse proximity is not explained by event or ideological identity alone. The semantic analysis further shows that the same geopolitical vocabulary acquires different evaluative orientations across conflicts and source environments. Finally, influencer-centered discussions exhibit higher toxicity on average, while the Iran corpora contain stronger hate-related signals that are also more persistent across replies. Together, these findings show that online war discourse is structured through the interaction of conflict context, media actors, audience communities, semantic framing, and conversational position.
2 Related Work
Political discourse on digital platforms is produced through connected flows of content, reaction, and recirculation rather than through isolated texts [7, 13]. Prior work has therefore examined how platform structures shape ideological exposure, polarization, and antagonistic interaction [1, 2]. YouTube is particularly relevant in this context because it combines persistent media objects, recommendation-driven navigation, and visible comment threads. Studies have investigated radicalization pathways, the circulation of problematic content, moderation effects, and contentious discussion within its channel and comment ecosystems [4, 9, 14, 21].
Political communication on YouTube is also shaped by the coexistence of institutional news organizations and creator-driven political commentary. Influencers increasingly act as intermediaries through whom audiences encounter and interpret political information, especially among younger users [18, 22]. Because news outlets and influencers differ in tone, audience relationships, and interpretive style, they may attract distinct forms of engagement. Chae and Lee [6], for example, identify differences in cross-cutting discussion around political vloggers and mainstream news outlets. These findings motivate our comparison of news- and influencer-centered communities across conflict contexts.
Computational approaches provide complementary tools for examining these communities. Relational and language-based embeddings can represent users and groups in latent spaces [11, 20], while WEAT-style association tests probe how concepts align with evaluative dimensions [5, 16]. POLAR extends this perspective through user-centered association measurement in embedding space [3]. Meanwhile, research on online hate speech emphasizes that hostile-language measurement is highly sensitive to context, dataset construction, and model generalization [12, 15, 17]. Building on these strands, we jointly compare community structure, semantic framing, and hostility across two conflicts and two types of media actors.
3 Data and Methods
3.1 YouTube Corpora
We collected data through the official YouTube Data API v3 [10] for two event-centered periods: the Afghanistan withdrawal and Taliban takeover from August to October 2021, and the Iran conflict escalation from February to April 2026. For each event, we separately considered institutional news organizations and political influencers, producing four subsets: Afghanistan–News, Afghanistan–Influencers, Iran–News, and Iran–Influencers.
Candidate channels were compiled from public rankings of U.S. YouTube news channels and English-language news and political creators. We retained channels that regularly published English-language political or news content and had at least one video matching the corresponding event filter. Videos were selected through title-based keyword searches, including terms such as Afghanistan, Taliban, Kabul, and withdrawal for the 2021 corpus, and Iran, Iranian, Tehran, and Ayatollah for the 2026 corpus.
To preserve channel diversity while maintaining comparable subset sizes, we sampled up to 100 videos per subset using a fixed random seed. We first selected one eligible video from each channel whenever available and then sampled uniformly from the remaining eligible videos until reaching the subset limit. For each selected video, we collected all accessible top-level comments and replies posted within the corresponding event window. The resulting dataset contains 340,383 comments from 212,968 unique commenters, as summarized in Table 1.
Table 1: Summary of the four event–source subsets. Comments include top-level comments and replies.
Subset | Channels | Comments | Commenters |
|---|---|---|---|
Afghanistan–Influencers | 15 | 131,602 | 79,146 |
Afghanistan–News | 9 | 71,690 | 46,346 |
Iran–Influencers | 22 | 87,998 | 56,770 |
Iran–News | 23 | 49,093 | 30,706 |
Total | – | 340,383 | 212,968 |
The dataset represents content visible through the API at collection time and therefore excludes deleted, moderated, private, disabled, or otherwise unavailable comments. User identifiers are used only for internal deduplication and structural analysis and are excluded from released artifacts. All collection and preprocessing code, complete channel lists, eligible and sampled video counts, sampling logs, keyword configurations, and supplementary materials are available in the accompanying repository.1
3.2 Community and Semantic Representations
To represent the discourse community surrounding each channel, we adapt POLAR [3] from user-level to channel-level representations. We define a channel–corpus unit as u = (c, g), where c is a channel and g is one of the four event–source subsets. Each unit receives a deterministic synthetic token tu, which is prepended to every comment associated with that unit. Thus, the same channel receives distinct representations in the Afghanistan and Iran corpora, allowing its surrounding discourse community to vary across conflicts.
We expand the tokenizer of bert-base-uncased with these synthetic tokens and fine-tune the model using masked-language-model adaptation following POLAR. The representation of each unit is the L2-normalized input embedding of its synthetic token. These vectors represent the linguistic environment produced by the comment community associated with a channel in a specific conflict, rather than the channel's uploaded content alone. Training parameters, preprocessing thresholds, balancing procedures, per-channel statistics, diagnostic metrics, and embedding files are provided in the repository.
To analyze the evaluative framing of geopolitical concepts, we train a single skip-gram Word2Vec model on the combined comment corpus. The model uses 300-dimensional vectors, a context window of five, a minimum frequency of five, and ten training epochs. A shared model is necessary because independently trained corpus embeddings would not be directly comparable.
We create corpus-specific versions of selected target terms by appending their event and source labels before training. For example, occurrences of trump in Afghanistan news comments and Iran news comments receive distinct tokens, while the remaining vocabulary is shared across all corpora. The targets include U.S. political actors and institutions, conflict regions, and war-related concepts such as trump, biden, america, iran, afghanistan, war, peace, and terrorism.
For target t in corpus k, we compute a single-target WEAT-style association score relative to positive and negative lexical attribute sets A and B:
Positive scores indicate greater proximity to the positive attribute set, whereas negative scores indicate greater proximity to the negative set. We estimate lexical sensitivity intervals by bootstrapping the attribute words. These intervals capture sensitivity to lexical-set composition rather than population uncertainty over comments or channels. For matched comparisons across conflicts, we test score differences within each source type and apply Benjamini–Hochberg false discovery rate correction. The complete attribute sets, normalization rules, bootstrap settings, and comparison outputs are available in the repository.
3.3 Hostility Proxies and Comparisons
We measure antagonistic discourse using two continuous classifier outputs rather than binary labels. Each comment is scored with unitary/toxic-bert, which provides our primary toxicity proxy, and cardiffnlp/twitter-roberta-base-hate-latest, which provides a secondary hate-related proxy. The resulting values lie in [0, 1] and are interpreted only as relative comparative signals. They should not be read as the probability that a comment is hateful or as the percentage of comments containing toxicity or hate speech.
We average these scores at three levels. First, we compare the four event–source subsets to examine differences between conflicts and between news- and influencer-centered discussion. Second, we separate top-level comments from replies to determine whether antagonistic language is concentrated in initial reactions or sustained through subsequent interaction. Third, we aggregate scores by channel to examine within-corpus variation among channel-centered communities.
Throughout the analysis, thread position, video, channel, event, and source-type metadata are preserved. Political orientation labels are used only as coarse qualitative descriptors for interpreting selected public channels. They are derived from public self-positioning or external outlet-level references, are never used as model inputs, and are not treated as ground-truth ideological measurements.
4 Results
In this section, we present results from the three components of our framework. First, we use POLAR-based channel embeddings to map channel-centered discourse communities into a shared geometric space and examine how they cluster across conflicts and source types. Second, we apply a WEAT-style lexical association analysis to identify how different communities align with salient geopolitical actors and concepts. Third, we compare the distribution of hostile and hate-related discourse across the four corpora using classifier-based proxies aggregated at multiple levels.
4.1 Channel-level discourse structure across conflicts
We begin with the global organization of the channel embedding space. Figure 1 shows a clear separation between the Afghanistan and Iran corpora. Afghanistan channels concentrate in one region of the projection, while Iran channels occupy another, with their centroids visibly far apart. Because each point represents the comment-centered discourse associated with a channel, rather than only the channel itself, this separation indicates that the public language circulating around the Iran conflict differs markedly from the language surrounding the Afghanistan withdrawal.
To check that this pattern is not only a projection artifact, we evaluated the separation in the original embedding space. The conflict split is clearly supported by these diagnostics: the cosine centroid-distance permutation test is significant (p = 0.001), and the conflict-based silhouette score is high (0.611). In contrast, source type shows little separation, with a silhouette score close to zero (0.003) and a non-significant permutation test. Thus, supporting the visual reading of Figure 1.
We further tested whether the t-SNE projection preserves the relevant structure of the learned embeddings. Across the tested perplexities and random seeds, trustworthiness remained high at both local neighborhood scales, around 0.95–0.96, and the projected conflict centroids remained significantly separated under permutation testing (p = 0.001).
The substantive interpretation of this separation is that the two conflicts generated distinct comment-centered discourse spaces. The Afghanistan corpus captures the final stage of a long war, whereas the Iran corpus corresponds to a new conflict escalation. The observed structure therefore suggests that news and influencer communities did not simply reproduce a generic discourse of war across both cases. Instead, the narratives surrounding each conflict appear to be organized differently in the comment space.
At the same time, the map is not divided into two perfectly isolated regions. Although the Afghanistan and Iran corpora are largely separated, a small set of Afghanistan channels and several Iran channels appear closer to one another near the center of the projection. This overlap is important because it shows that the geometry is not determined only by event labels. If conflict context alone explained the embedding space, we would expect the two corpora to remain fully apart. Instead, the partial overlap suggests that some channels from different events converge toward similar discourse patterns.
This convergence becomes especially informative when we examine which channels occupy the central region. On the Afghanistan side, the two influencer channels closest to this area are Ben Shapiro and Allie Beth Stuckey, both publicly associated with conservative political commentary. On the Iran side, nearby channels are mostly news outlets generally associated with more left-leaning or progressive editorial orientations, such as Vox, Democracy Now!, Business Insider, and VICE News. Thus, the overlap is not formed by actors with the same political identity. Rather, it brings together actors from different ideological positions whose surrounding comment communities nevertheless become close in the discourse space.
This pattern suggests that proximity in the channel embedding space does not simply reflect ideological alignment. Instead, it may capture shared styles of conflict commentary, including more reactive, emotional, or interpretive forms of discussion. Together with the embedding-space diagnostics, Figure 1 therefore supports a nuanced interpretation: YouTube war discourse is strongly organized by conflict context, but it also contains zones of cross-context convergence among politically different actors.
4.2 Semantic Associations
We next examine whether the separation observed in the channel-embedding space is also reflected in the semantic associations of key geopolitical and conflict-related terms. Figure 2 reports single-target WEAT-style association scores for trump, america, war, biden, iran, afghanistan, peace, and terrorism across the four corpora. Positive values indicate stronger proximity to the positive attribute set, while negative values indicate stronger proximity to the negative attribute set. Because the reported intervals are based on attribute resampling, we interpret them as lexical sensitivity intervals rather than full population-level confidence intervals.
The clearest descriptive pattern is that peace receives positive point estimates across all four corpora, while terrorism receives negative point estimates across all four corpora. We use these terms mainly as a sanity check for the measure, since they are strongly valenced and therefore provide a simple test of whether the evaluative axis behaves in the expected direction.
Other conflict-related terms show more contextual variation. Iran leans negative across all four corpora, suggesting that it tends to appear in evaluative contexts related to threat, escalation, or criticism. By contrast, Afghanistan receives positive point estimates in all four corpora, indicating that discussions around Afghanistan are comparatively less tied to negative evaluative vocabulary and may be more associated with aftermath, withdrawal, or humanitarian framing. These patterns are consistent across source types, although their intervals remain wide.
The U.S. political terms show the largest descriptive shifts across conflict contexts. In the news-centered corpora, trump moves from near-neutral in 2021 news comments to clearly negative in Iran-related news comments in 2026. A similar but weaker pattern appears in influencer-centered comments, where trump is positive in the Afghanistan corpus and closer to neutral or slightly negative in the Iran corpus. America follows a related pattern: it is positive in Afghanistan-related discussions, especially in influencer comments, but becomes less positive or slightly negative in Iran-related news comments. These shifts suggest that U.S.-related political vocabulary is evaluated differently across conflict contexts.
By comparison, biden and war show weaker or more uneven patterns. Biden remains close to neutral across most corpora, with only a slight negative tendency in Iran-related news comments. War is more negative in Iran-related news comments than in Afghanistan-related news comments, but it remains close to neutral in the influencer corpora. Overall, the WEAT-style analysis suggests that YouTube war discourse is not organized by topic alone: the same geopolitical vocabulary can acquire different evaluative orientations depending on conflict context and source environment.
4.3 Hate Speech and Toxicity
We begin this analysis by comparing the levels of hostility and hate-related language across the two events and across the two source types, namely news channels and influencers. Importantly, the values reported in this subsection are based on mean aggregated proxy scores produced by our models. They should therefore not be interpreted as the percentage of comments that contain hate speech or toxicity. Instead, they provide a relative measure that allows us to compare the four datasets with one another and to examine how antagonistic discourse varies across time periods, media actors, and comment structures.
In Figure 3, we compare the average toxicity and hate-related proxy scores across the four datasets. The first pattern that stands out is toxicity: discussions in influencer-centered videos are consistently more toxic than those in news-centered videos, with the Iran collections showing slightly higher values than the Afghanistan ones in both source types. This result is not entirely surprising. Influencer content is often more personalized, emotional, and openly partisan, which can encourage more confrontational styles of engagement in the comment section. News channels, by contrast, tend to structure discussion around more institutional and report-driven content, which may still be contentious but usually produces somewhat less aggressive comment spaces on average.
A second and especially notable pattern appears in the hate-related proxy. Here, both Iran datasets score clearly higher than their Afghanistan counterparts, suggesting that hate-related language became more salient in the 2026 conflict context than in the 2021 one. This was broadly in line with our expectations. The Iran conflict emerged in a period of heightened global polarization, intensified geopolitical tensions, and highly reactive online discourse, all of which may have contributed to a sharper escalation in antagonistic language [8, 25]. Taken together, these results suggest that the difference between 2021 and 2026 is not only one of overall toxicity, but also one of how strongly conflict-related discussions become associated with more explicitly hostile and hate-related forms of expression.
Figure 4 adds an important structural layer to the previous comparison by separating top-level comments from replies. A clear pattern emerges in both panels: across all four datasets, top-level comments are consistently more toxic and more hate-related, on average, than replies. This suggests that antagonistic discourse is introduced more strongly in the initial reaction to a video than in the subsequent back-and-forth of the thread. In the toxicity proxy, the gap is especially large for Afghanistan–Influencers, while it is smaller but still consistent in Afghanistan–News, Iran–News, and Iran–Influencers. In other words, toxicity tends to be front-loaded: the opening comments are the most confrontational part of the discussion, while replies remain antagonistic but somewhat less intense.
The hate-speech proxy shows the same overall tendency, but with a more nuanced temporal pattern. In the Afghanistan datasets, the drop from top-level comments to replies is relatively pronounced, particularly for Afghanistan–News and Afghanistan–Influencers. In the Iran datasets, however, this gap becomes smaller, especially for Iran–Influencers, where top-level comments and replies are very close. This is a notable result: while top-level comments still carry the highest mean scores, hate-related language in the Iran discussion appears to be more evenly sustained throughout the thread rather than remaining concentrated only in the initial reaction. Taken together, these findings suggest that the 2026 Iran conflict not only produced higher overall levels of antagonistic discourse, but also a more persistent spread of that discourse into conversational replies, particularly in influencer-centered spaces.
Figure 5 extends the dataset-level comparison by showing that each corpus contains a non-trivial internal spread at the channel level. Rather than forming a single compact cluster, the channels occupy different positions in the joint space defined by the toxicity proxy and the hate-speech proxy, indicating that antagonistic discourse is unevenly distributed within each collection. Among the four panels, Afghanistan–News appears comparatively tighter and more moderate, whereas Iran–Influencers is the most dispersed, suggesting a more heterogeneous and polarized discussion environment. Iran–News also occupies a visibly higher region on the hate-related axis than Afghanistan–News, which reinforces the earlier aggregate result that the 2026 Iran conflict was associated with more strongly antagonistic discourse than the 2021 Afghanistan case. In this sense, the difference between the two events is not only a corpus-wide average effect, but also a channel-level effect that appears across multiple actors within the same media type.
The annotated channels help interpret this geometry politically. In Afghanistan–Influencers, the highlighted channels are broadly aligned with the political right, which makes that panel relatively coherent in ideological terms and is consistent with a comment space shaped by conservative and security-oriented framing. Afghanistan–News is more mixed: its highlighted channels combine center or lean-left outlets with a channel from the Fox ecosystem, which fits the more compact and less polarized appearance of that panel. Iran–News is especially notable because the highlighted channels are broadly left or lean-left, suggesting that the higher hate-related values in this panel are not simply a right-wing phenomenon, but may also reflect the intensity of oppositional or highly critical discourse around the conflict. Finally, Iran–Influencers is the most ideologically mixed of the four, combining a clearly left-leaning channel with conservative ones. This suggests that the 2026 Iran conflict produced antagonistic discourse across partisan camps, with especially strong channel-level variation emerging in highly opinionated influencer spaces rather than being confined to a single ideological side.
5 Conclusion and Ethical Considerations
This paper compared YouTube discourse surrounding the Afghanistan withdrawal in 2021 and the Iran conflict escalation in 2026 across news and influencer channels. The results show that conflict context is the main organizing dimension of channel-centered discourse communities, while source type produces little global separation. The semantic analysis further indicates that geopolitical terms acquire different evaluative orientations across conflicts and media environments. Finally, influencer-centered discussions exhibit higher toxicity, whereas the Iran corpora contain stronger and more persistent hate-related signals, including across replies. Together, these findings show that online war discourse is shaped by the interaction of event context, media actors, audience communities, and conversational structure.
We analyze only publicly available data collected through the official YouTube API. We do not report usernames, commenter identifiers, comment identifiers, or verbatim comments, and released artifacts exclude personal identifiers. Toxicity and hate-related scores are interpreted as comparative proxies rather than definitive labels, since automated classifiers may produce errors in politically sensitive or context-dependent language. The dataset also excludes comments that were deleted, moderated, private, disabled, or otherwise unavailable at collection time. Political orientation labels are used only as coarse qualitative descriptors and never as model inputs or ground-truth ideological measurements.
Source
Imported from ACM’s structured HTML source. ACM Reference Format: Arthur Buzelin, Pedro Augusto Torres Bento, Yan Aquino Amorim, Arthur Rodrigues Chagas, Virgilio Almeida, Gisele Lobo Pappa, and Wagner Meira Jr.. 2026. YouTube Discourse During the Iran Conflict and the Afghanistan Withdrawal. In 37th ACM Conference on Hypertext (HT '26), September 14--18, 2026, London, United Kingdom. ACM, New York, NY, USA 7 Pages. https://doi.org/10.1145/3800935.3830874
References
[1] Christopher A. Bail, Lisa P. Argyle, Taylor W. Brown, John P. Bumpus, Haohan Chen, M. B. Fallin Hunzaker, Jaemin Lee, Marcus Mann, Friedolin Merhout, and Alexander Volfovsky. 2018. Exposure to opposing views on social media can increase political polarization. Proceedings of the National Academy of Sciences 115, 37 (2018), 9216–9221. https://doi.org/10.1073/pnas.1804840115
[2] Eytan Bakshy, Solomon Messing, and Lada A. Adamic. 2015. Exposure to ideologically diverse news and opinion on Facebook. Science 348, 6239 (2015), 1130–1132. https://doi.org/10.1126/science.aaa1160
[3] Pedro Bento, Arthur Buzelin, Arthur Chagas, Yan Aquino, Victoria Estanislau, Samira Malaquias, Pedro Robles Dutenhefner, Gisele L. Pappa, Virgilio Almeida, and Wagner Meira Jr.2026. POLAR: A Per-User Association Test in Embedding Space. In Proceedings of the 20th International AAAI Conference on Web and Social Media(ICWSM 2026).
[4] Cody Buntain, Richard Bonneau, Jonathan Nagler, and Joshua A. Tucker. 2021. YouTube Recommendations and Effects on Sharing Across Online Social Platforms. Proc. ACM Hum. Comput. Interact. 5, CSCW1 (2021), 11:1–11:26. https://doi.org/10.1145/3449085
[5] Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science 356, 6334 (2017), 183–186. https://doi.org/10.1126/science.aal4230
[6] Seung Woo Chae and Sung Hyun Lee. 2024. Where do cross-cutting discussions happen?: Identifying cross-cutting comments on YouTube videos of political vloggers and mainstream news outlets. PLOS ONE 19, 5 (2024), e0302030. https://doi.org/10.1371/journal.pone.0302030
[7] danah boyd. 2010. Social Network Sites as Networked Publics: Affordances, Dynamics, and Implications. In A Networked Self: Identity, Community, and Culture on Social Network Sites, Zizi Papacharissi (Ed.). Routledge, London, UK, 39–58. https://doi.org/10.4324/9780203876527-8
[8] Edelman Trust Institute. 2025. 2025 Edelman Trust Barometer: Trust and the Crisis of Grievance. https://www.edelman.com/trust/2025/trust-barometer
[9] Gabriel Luís Santos Freire, Tales Panoutsos, Lucas Perez Santos, Fabrício Benevenuto, and Flavio Figueiredo. 2022. Understanding Effects of Moderation and Migration on Online Video Sharing Platforms. In Proceedings of the 33rd ACM Conference on Hypertext and Social Media. ACM, 220–224. https://doi.org/10.1145/3511095.3536377
[10] Google Developers. 2026. YouTube Data API v3. https://developers.google.com/youtube/v3. Accessed: 2026-04-13.
[11] Aditya Grover and Jure Leskovec. 2016. node2vec: Scalable Feature Learning for Networks. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery, 855–864. https://doi.org/10.1145/2939672.2939754
[12] Samuel S. Guimarães, Gabriel Kakizaki, Philipe F. Melo, Márcio Silva, Fabricio Murai, Julio C. S. Reis, and Fabrício Benevenuto. 2023. Anatomy of Hate Speech Datasets: Composition Analysis and Cross-dataset Classification. In Proceedings of the 34th ACM Conference on Hypertext and Social Media (HT 2023), Rome, Italy, September 4-8, 2023. ACM, 33:1–33:11. https://doi.org/10.1145/3603163.3609158
[13] George P. Landow. 2006. Hypertext 3.0: Critical Theory and New Media in an Era of Globalization (3 ed.). Johns Hopkins University Press, Baltimore, MD, USA.
[14] Marcelo Sartori Locatelli, Josemar Alves Caetano, Wagner Meira Jr., and Virgílio A. F. Almeida. 2022. Characterizing Vaccination Movements on YouTube in the United States and Brazil. In HT ’22: 33rd ACM Conference on Hypertext and Social Media, Barcelona, Spain, 28 June 2022- 1 July 2022. ACM, 80–90. https://doi.org/10.1145/3511095.3531283
[15] Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021. HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 14867–14875. https://doi.org/10.1609/aaai.v35i17.17745
[16] Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019. On Measuring Social Biases in Sentence Encoders. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Association for Computational Linguistics, Minneapolis, Minnesota, 622–628. https://doi.org/10.18653/v1/N19-1063
[17] Mainack Mondal, Leandro Araújo Silva, and Fabrício Benevenuto. 2017. A Measurement Study of Hate Speech in Social Media. In Proceedings of the 28th ACM Conference on Hypertext and Social Media (HT 2017), Prague, Czech Republic, July 4-7, 2017. ACM, 85–94. https://doi.org/10.1145/3078714.3078723
[18] Nic Newman, Amy Ross Arguedas, Mitali Mukherjee, and Richard Fletcher. 2025. Mapping News Creators and Influencers in Social and Video Networks. Technical Report. Reuters Institute for the Study of Journalism, University of Oxford. https://doi.org/10.60625/risj-44pf-1k13
[19] OECD. 2025. How's Life for Children in the Digital Age?Technical Report. Organisation for Economic Co-operation and Development. https://www.oecd.org/content/dam/oecd/en/publications/reports/2025/05/how-s-life-for-children-in-the-digital-agec4a22655/0854b900-en.pdf Accessed: 2026-04-27.
[20] Bryan Perozzi, Rami Al-Rfou, and Steven Skiena. 2014. DeepWalk: Online Learning of Social Representations. In Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. Association for Computing Machinery, 701–710. https://doi.org/10.1145/2623330.2623732
[21] Manoel Horta Ribeiro, Raphael Ottoni, Robert West, Virgílio A. F. Almeida, and Wagner Meira Jr.2020. Auditing radicalization pathways on YouTube. In FAT ’20: Conference on Fairness, Accountability, and Transparency, Barcelona, Spain, January 27-30, 2020, Mireille Hildebrandt, Carlos Castillo, L. Elisa Celis, Salvatore Ruggieri, Linnet Taylor, and Gabriela Zanfir-Fortuna (Eds.). ACM, 131–141. https://doi.org/10.1145/3351095.3372879
[22] Galen Stocking, Luxuan Wang, Michael Lipka, Katerina Eva Matsa, Regina Widjaya, Emily Tomasik, and Jacob Liedke. 2024. America's News Influencers. Technical Report. Pew Research Center. https://www.pewresearch.org/journalism/2024/11/18/americas-news-influencers/
[23] UNESCO. 2024. UNESCO dedicates the International Day of Education 2024 to countering hate speech. https://www.unesco.org/en/articles/unesco-dedicates-international-day-education-2024-countering-hate-speech. Accessed: 2026-04-27.
[24] United Nations. 2023. Information Integrity on Digital Platforms. Technical Report. United Nations. https://brasil.un.org/sites/default/files/2023-06/our-common-agenda-policy-brief-information-integrity-en.pdf Policy Brief No. 8, Accessed: 2026-04-27.
[25] World Economic Forum. 2026. The Global Risks Report 2026. https://www.weforum.org/publications/global-risks-report-2026/
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime