Link-Traced Amplification: How Telegram Redistributes YouTube Political Content at ScaleTelegram is a lightly moderated platform hosting, among others, fringe and politically extreme communities, many of which migrated after being deplatformed from mainstream social media.

Abstract

Telegram is a lightly moderated platform hosting, among others, fringe and politically extreme communities, many of which migrated after being deplatformed from mainstream social media. While prior research has mainly analyzed political debate on Telegram through in-platform text, less attention has been paid to how external media content flows into and circulates within this ecosystem. We study the circulation of YouTube videos shared in 43,000 public Telegram chats surrounding the 2024 U.S. presidential election, analyzing 686,625 English-language videos as traces of cross-platform connectivity. Using topic modeling, supervised classification, and multi-dimensional toxicity measures, we characterize which narratives are amplified, how they diffuse by analyzing reach, recirculation, persistence, and transfer time, and whether toxicity is related to amplification. We find that Telegram functions as an agenda-redistribution layer for YouTube political content. Political videos diffuse in bursty, short-lived patterns that track high-salience events. Toxicity is only weakly associated with reach; instead, more toxic content tends to persist longer within narrower thematic circuits, reinforcing segmented information environments.

1 Introduction

Telegram has emerged as a relevant infrastructure for political communication in polarized and high-stakes contexts. Its architecture, with large public channels and comparatively low moderation, has made it attractive to fringe and far-right communities, including groups displaced from mainstream platforms [48, 51, 53]. Beyond hosting extremist and conspiratorial narratives, Telegram has been associated with the dissemination of misinformation and coordinated political mobilization during electoral cycles [2, 24, 52]. Recent large-scale datasets document the platform's scale and structural complexity, spanning tens of thousands of public chats and hundreds of millions of messages [5, 14, 22, 30]. Prior work has examined platform usage, ideological segmentation and information propagation within Telegram communities and across topics [27, 43], while studies of cross-platform coordinated activity highlight the importance of shared URLs in contemporary information operations [12, 50, 54]. These findings situate Telegram within the broader hybrid media system, where political narratives travel across interconnected platforms rather than remaining confined to a single environment [8, 9]. However, most existing work focuses on in-platform textual discourse or coordination dynamics [21, 40], leaving unexplored how external media objects, particularly political videos, flow into Telegram and are redistributed within its public-sphere infrastructure.

This paper addresses this gap by characterizing YouTube links shared in public Telegram channels at scale. We analyze English-language videos disseminated during 2024, a year marked by intense electoral and geopolitical salience, notably in the U.S.. Shared URLs occupy a structurally distinct position in this ecosystem: unlike textual posts, they are explicit linking acts that bridge two platform environments, transporting not only content but also the interpretive context in which it is received. By foregrounding this hypertextual layer, we conceptualize Telegram not merely as a discussion space but as a redistribution infrastructure that selectively amplifies political video content produced elsewhere. This link-centric perspective distinguishes our approach from studies that model platform discourse through in-platform text alone.

We investigate three research questions (RQs):

RQ1: Which topics emerge from YouTube videos shared within U.S. political Telegram chats during 2024, and how does toxicity vary across them?

RQ2: How do these videos diffuse within Telegram in terms of reach, recirculation, and persistence, and how does toxicity relate to amplification?

RQ3: Are sharing dynamics primarily event-driven or persistent over time, and how do these temporal regimes differ between political and non-political topics?

To answer these questions, we extracted 1,911,054 YouTube links from a previously collected Telegram dataset [5], composed of one billion Telegram messages across 43,000 public U.S. political chats 1. We retrieve each video's title and description via the YouTube Data API [19].We then apply BERTopic [20] with supervised classification to identify thematic categories and use the Perspective API [31] to measure multiple toxicity dimensions, enabling large-scale comparisons of diffusion patterns across topics.

Our findings indicate that Telegram functions as an agenda-redistribution mechanism: shared videos closely track high-visibility events and media trends, yet exhibit meaningful asymmetries across thematic categories. Political content tends to be more toxic than non-political content, with macro-topics related to War and Crime concentrating the highest toxicity levels. Diffusion also differs qualitatively: political videos, especially during election periods, when rapid events and emerging debates continuously reshape attention, tend to generate sharp peaks in dissemination within short time windows, whereas non-political videos, covering topics such as Entertainment and Technology, accumulate attention more gradually and over longer horizons. Over time, themes tied to U.S. politics gain prominence and drive event-based spikes in activity, while average toxicity remains comparatively stable, suggesting that amplification is structured primarily by topics rather than isolated items.

In sum, this study makes three main contributions: (i) it operationalizes cross-platform circulation by treating shared links as observable traces of how content moves between platforms; (ii) it introduces a pipeline for large-scale data enrichment, topic labeling, and toxicity measurement, with the implementation publicly available online2; and (iii) it provides an empirical characterization of YouTube → Telegram diffusion patterns, clarifying when toxicity relates to persistence rather than reach.

2 Related Work

Research on Telegram political communication spans extremist mobilization [28], large-scale discourse mapping [40], and coordinated information operations [12]. Studies document the migration of fringe actors from mainstream platforms and the consolidation of ideological networks within public Telegram channels [6, 49, 51, 53]. The increasing availability of large public datasets [5, 30] has enabled ecosystem-level analyses spanning tens of thousands of chats and hundreds of millions of messages. Building on these resources, prior work has mapped topic structure, participation triggers, community organization, and toxic discourse patterns at scale [1, 39, 43, 52]. Yet, these studies primarily model internal message-level dynamics.

A parallel strand investigates information diffusion and coordinated amplification within and across platforms. Forwarded messages and embedded URLs play a central role in cross-chat propagation [27]. Studies of coordinated behavior demonstrate how shared URLs and synchronized posting patterns reveal organized amplification campaigns, including cross-platform coordination during elections [12, 54]. Research on messaging platforms further shows how external media objects circulate within political chats, including YouTube videos shared in electoral contexts [25, 45, 46], while identifying thematic and actor-level patterns of cross-platform amplification [11]. Yet these analyses typically focus on specific communities or limited contexts and do not jointly model thematic structure, diffusion regimes, and cross-platform amplification at scale.

A complementary tradition treats hyperlinks not merely as technical pointers but as meaningful communicative acts that structure the visibility of information across networked environments [4]. Cross-platform studies have operationalized this view by tracing shared URLs and media objects across distinct online ecosystems: fringe communities have been shown to act as injection points that push content, whether news URLs or image memes, into broader information flows, with measurable asymmetries in timing and reach across platforms [55, 56]. Prior work has further shown that cross-platform URL sharing can constitute a structural bypass of content governance, with harmful YouTube videos continuing to circulate on Twitter after moderation [16]. We adopt this link-centric stance, treating each shared YouTube URL as an observable trace of cross-platform bridging rather than a mere metadata field.

Finally, automated detection of harmful and toxic speech has been widely studied in political and social media settings [15]. Prior work has analyzed YouTube content using topic, engagement, and toxicity metrics in specific thematic contexts [33], while recent audits have examined how the platform handles politically sensitive and conspiratorial content, highlighting governance gaps that motivate the study of such material's migration to less-moderated environments [18]. Toxicity measures derived from tools such as the Perspective API have been used to characterize ideological ecosystems on YouTube [35], although methodological work cautions against equating such scores with normative constructs of incivility [17].

While prior research provides large-scale mappings of Telegram discourse, evidence of URL-mediated coordination, and computational frameworks for toxicity measurement, we lack a link-centric characterization of which YouTube videos are amplified within U.S. political Telegram chats, how diffusion unfolds in terms of reach and persistence, and how toxicity varies across thematic regimes along the YouTube → Telegram pathway. We address this gap by integrating topic modeling, supervised labeling, toxicity measurement, and diffusion analysis to quantify cross-platform amplification at corpus scale.

3 Experimental Methodology

Our methodological pipeline studies the cross-platform circulation of political discourse by treating shared YouTube URLs on Telegram as observable hypertextual traces. Starting from a large-scale corpus of public Telegram messages, we (i) construct a link-centric dataset that captures when and where YouTube videos are posted across public chats; (ii) enrich each unique video by collecting public metadata via the YouTube Data API; (iii) induce a topic taxonomy from title–description text using BERTopic and extend it to the full corpus through supervised macro-topic classification; and (iv) quantify multiple dimensions of toxicity with the Perspective API to relate topical structure and temporal diffusion patterns to the prevalence of hostile and identity-targeting language.

3.1 Dataset

We build our analysis on the large-scale public Telegram dataset released by Blas et al. [5], which documents political discourse surrounding the 2024 U.S. presidential election. The collection comprises approximately 1.03 billion messages spanning 43k public Telegram chats and covering the period from December 2023 to January 2025. For each message, the dataset provides several metadata fields. In this study, we use only the message identifier, timestamp, textual content, and chat identifier, which are sufficient to identify YouTube URLs and analyze their temporal distribution and reach across public Telegram chats.

3.1.1 URL-derived dataset construction. Figure 1 depicts the monthly dynamics of link sharing (counting all occurrences, including repeated postings of the same URL). We observe a sharp peak during the U.S. election month, with more than 100 million URLs circulating across the analyzed chats, suggesting that Telegram operates as a redistribution infrastructure during the period of heightened political attention.

Figure 1: Number of URLs shared per month (all occurrences).

3.1.2 Filtering by social-platform domains. A manual inspection of the extracted URLs revealed substantial noise, with many domains falling outside the scope of this study (e.g., cryptocurrency, NFTs, promotional pages). We therefore applied a domain-based filter, retaining only links associated with major social platforms relevant to political information circulation. This filtering step results in 6,507,630 unique links, distributed as shown in Table 1.

The most prevalent domains are X and YouTube, consistent with recent findings: [43] reports that YouTube, X, and Instagram are among the most frequently shared external domains on Telegram, with YouTube appearing across topics. Given this centrality and the availability of public metadata via official APIs, we focus on the YouTube → Telegram interaction.

Table 1: Link Distribution by Platform (after domain filtering).

Platform

Count

Retained domains

X (Twitter)

3,811,209

twitter.com, x.com, t.co

YouTube

1,911,054

youtube.com, youtu.be

Instagram

300,291

instagram.com

Rumble

175,970

rumble.com

TikTok

143,217

tiktok.com

Facebook

108,586

facebook.com, fb.com, fb.watch

VK

53,960

vk.com, vkontakte.ru

TruthSocial

3,343

truthsocial.com

3.1.3 URL normalization and video ID selection. The initial extraction identified 1,911,054 YouTube URLs. We then retain only URLs corresponding to consumable content units: videos, Shorts, and live streams. We remove links pointing to YouTube channels, playlists, and structural pages, as these do not constitute content items with directly comparable metadata and are not aligned with our analytical objective. After this step, we extract the corresponding identifiers (videoId) and obtain 1,397,385 distinct video IDs, since multiple URL variants (e.g., parameterizations, shorteners, and format differences) can resolve to the same video.

3.1.4 Enrichment via the YouTube Data API. For each valid videoId, we enriched the dataset (as of January 2026) using the YouTube Data API v3 [19] to retrieve public metadata required for topic modeling, toxicity inference, and subsequent analyses. Specifically, we retrieved two sets of metadata: (i) snippet: title, description, channel, and publication date; (ii) statistics: views, likes, and comments.

This procedure yields 1,183,816 valid video IDs with complete metadata. The difference relative to the total number of extracted IDs reflects unavailability at the time of collection (e.g., removed or private videos, region- or age-based restrictions, or transient errors), which we log and exclude from downstream analyses.

3.1.5 Final language filtering. The enriched corpus spans multiple languages. To ensure linguistic cohesion and improve the performance of downstream Natural Language Processing (NLP) techniques, we perform language detection on the concatenated text (title + description) using langdetect [37]. For each item, the detector returns the most likely language and an associated probability. We retained only items for which English was the language with the highest predicted probability, consistent with the US-focused scope of the dataset. After filtering, the final dataset totals 686,625 videos.

3.1.6 Data Pre-processing. We apply a fixed document-level preprocessing pipeline including lowercasing, lemmatization with spaCy [26], removal of English stopwords (using the standard NLTK list [3]), diacritic stripping, removal of social-media markers, URLs, punctuation, and digits, normalization of elongated tokens, and filtering of tokens with length ≤ 3.

3.2 Topic Modeling

To identify the predominant themes in YouTube content shared on Telegram, we perform topic modeling over the text obtained by concatenating each video's title and description. We rely on BERTopic [20], a state-of-the-art transformer-based approach that can induce semantically coherent topics even in large-scale corpora with highly heterogeneous language. As illustrated in Figure 2, our end-to-end mining workflow comprises five main stages. Given the scale of the dataset, directly modeling the entire corpus would be computationally prohibitive. We therefore adopt a structured sampling strategy that preserves topical diversity while ensuring tractability, enabling a faithful characterization of the principal themes expressed in the text.

Figure 2: Topic modeling pipeline.

Hyperparameter Selection. We next sought an adequate BERTopic configuration for our dataset by performing a systematic hyperparameter exploration on a 1% sample of the corpus (6,735 videos), using concatenated title+description as the textual unit. The parameters and candidate values are listed in Table 2, yielding 648 possible configurations. We evaluate each configuration using three complementary metrics, namely topic coherence [36], topic diversity [13], and silhouette score [47], as well as the average of these three measures. Because no single metric fully captures topic quality [29], this multi-metric evaluation allows us to balance semantic interpretability, topical coverage, and embedding-space separability.

Table 2: Hyperparameters and candidate values.

Parameters

Possibilities

umapneighbors

10, 20

umapcomponents

5, 10

umapmindist

0.0, 0.1, 0.5

hdbscanmincluster

50, 100, 150

hdbscanminsamples

5, 10, 100

hdbscanclustereps

0.1, 0.5

embeddingmodels

all-MiniLM-L6-v2, all-distilroberta-v1,


paraphrase-MiniLM-L6-v2

We select, for each metric, the top-20 configurations, as the differences among the best-performing settings are marginal, indicating near-equivalent performance under the automatic criteria. We then apply BERTopic to a larger 10% sample of the full dataset (67,359 videos), running one model per selected configuration, for a total of 80 topic-modeling runs. We further assess each run through qualitative manual inspection of the resulting topics to identify the most interpretable and coherent solution. Among the 80 candidates, we select the configuration achieving coherence = 0.635, diversity = 0.996, silhouette = 0.824, and mean = 0.818, with parameters: umapneighbors= 10, umapcomponents= 10, umapmindist= 0.0, hdbscanmincluster= 100, hdbscanminsamples= 100, hdbscanclustereps= 0.1, and embeddingmodels= paraphrase-MiniLM-L6-v2.3

Macro-topics Construction. To support efficient yet representative interpretation, we adopt a multi-evidence summarization strategy. For each topic, we extract: (i) the ten most representative keywords, (ii) the three most representative documents according to BERTopic, and (iii) ten randomly sampled videos. We then use ChatGPT-5.4 Thinking model [38] to produce a consolidated topic summary. Based on this output, we manually review and validate the topic labels and descriptions to ensure accurate interpretation. To enable generalization, we perform a manual semantic aggregation step, reorganizing topics into a smaller number of conceptually cohesive broader categories, which we refer to as macro-topics. The resulting macro-topics are: (1) U.S. Politics; (2) Foreign Politics; (3) Crime, Security, and Justice; (4) War; (5) Religion; (6) Technology; (7) Health; (8) Environment; (9) Entertainment; (10) Hobbies; and (11) Other (aggregating outliers and topics outside the analytical scope, such as cryptocurrencies).

Although the dataset was originally collected using politically-oriented seed terms in the context of the 2024 U.S. elections [5], which may introduce sampling biases in both topical coverage and temporal dynamics, given the heightened activity and rapid influx of information typical of election, prior work has shown that the snowball-based collection process naturally expands beyond strictly political groups, capturing a substantial amount of non-political content [39]. Therefore, while our findings should be interpreted in light of this potential bias, the dataset still enables a meaningful comparison between political and non-political content.

Thus, we consolidated these macro-topics into two analytical sets: political (U.S. Politics, Foreign Politics, War, and Crime/Security/Justice) and non-political (the remaining ones). A macro-topic was classified as political when its representative content predominantly addressed governmental institutions, political actors, public policies, electoral processes, or geopolitical conflict. Categories whose themes did not primarily concern these domains were classified as non-political, even when occasional intersections with political discourse occurred.

3.3 Classifying the Remaining Videos

With macro-topics manually assigned to the 10% sample, we train a supervised classifier to automatically label the remainder of the dataset. We represent each document as a dense semantic vector using paraphrase-MiniLM-L6-v2, the embedding model selected during BERTopic tuning, and employ Random Forest [7], a widely used classifier in text categorization due to its strong performance in high-dimensional settings [32, 41]. We use stratified 5-fold cross-validation (Stratified K-Fold), grid search for hyperparameter optimization, and Macro F1 as the primary evaluation metric. The best hyperparameters are: classweight=balanced, maxdepth=None, minsamplessplit=5, and nestimators=200. The final obtained Macro F1 score is 0.753 with standard deviation 0.005, indicating stable behavior across folds. We then apply the trained model to label the full dataset, enabling temporal and thematic analyses at scale.

3.3.1 Macro-topic summarization. After labelling the full dataset into macro-topics, we construct interpretable semantic descriptions for each category. Since macro-topics aggregate multiple BERTopic topics, we no longer rely on topic-model artifacts (e.g., keywords and representative documents), but instead summarize representative documents directly. For each macro-topic, we sample up to 20,000 documents and embed their title+description texts using paraphrase-MiniLM-L6-v2. We then compute the semantic centroid (mean embedding), which corresponds to the region of highest semantic density, and rank documents by inner-product similarity. The 20 closest documents are selected as macro-topic exemplars, acting as representatives of their central structure, and summarized by ChatGPT-5.4 Thinking, producing a concise, human-interpretable description, which is subsequently manually validated. To ensure consistency, neutrality, and reproducibility, we use the following standardized prompt:

Prompt

The input consists of documents comprising the titles and descriptions of YouTube videos, grouped into macro-topics.

Input: top20docs: The 20 most representative documents for the macrotopic.

Instructions: Carefully analyze the provided content to infer the main theme of the macrotopic grouping; Focus on identifying recurring patterns, themes, or subjects in the transcripts; Do not mention or reference specific videos, speakers, or quotes; Use clear, neutral, and concise language; Do not invent information not present in the inputs.

I will send the information in a file. Considering the information provided, prepare the summary following this format:

Macrotopic Title: <macrotopic title>; Summary: <macrotopic summary>

The answer must be strictly in English.


3.4 Video Toxicity Estimation

To quantify toxicity associated with the videos, we assume that, in the absence of full transcripts, the publicly available textual fields (title and description) provide a scalable proxy for the tone and framing of the underlying content. Accordingly, these textual fields were used as input for toxicity inference.

Since there is no single universally accepted model for toxicity detection, and different tools may capture distinct aspects of harmful language depending on the domain [10, 42], platform, and textual cues available, we experiment with two widely used toxicity detection systems. We first apply Detoxify [23], a BERT-based model executed in an offline setting, which outputs scores in [0,1] across multiple toxicity-related dimensions (including overall toxicity, severe toxicity, threat, insult, and identity attack). In parallel, we also evaluate toxicity using Google's Perspective API [31], requesting the same dimensions.

For model comparison, we focus on the overall toxicity score and compute the absolute difference between Detoxify and Perspective outputs for each video. We identify cases of substantial divergence and manually inspect a sample of the most discrepant items to assess the nature of disagreement. This inspection reveals two recurring patterns. First, in our dataset, Detoxify tends to generate context-sensitive false positives, assigning elevated toxicity scores to non-violent or metaphorical uses of terms such as “war” or “battle” (e.g., theological or gaming contexts). Second, Detoxify appears less sensitive to implicit hostility expressed without explicit profanity, whereas Perspective more consistently captures tone-dependent threats and antagonistic framing.

Based on qualitative auditing of strongly divergent cases and the observed contextual robustness of Perspective in our dataset, we adopt the Perspective API as the primary method for toxicity measurement in the remainder of the study. We therefore apply it to the full corpus of videos, retaining the same toxicity dimensions for downstream analyses.

4 Results

Our analysis unfolds along three dimensions that mirror our research questions: how shared YouTube links organize into thematic domains, how they propagate across Telegram's public chats, and how these hypertextual circulation patterns evolve over time.

4.1 Thematic Landscape and Toxicity (RQ1)

4.1.1 What is Shared on Telegram: Macro-topics. To understand what types of content are redistributed through YouTube → Telegram’s hyperlink layer, we begin by examining how content from shared YouTube links clusters into macro-topics. As described in Section 3.3.1, we produced concise semantic summaries for each macro-topic. For space reasons, we report here only the summaries of the four political macro-topics, which constitute the core analytical focus of the study, while the others are largely self-descriptive.

Figure 3: Number of shared YouTube videos per macro-topic.
Figure 4: Distribution of video overall toxicity by macro-topic.

U.S. Politics: centers on contemporary United States political dynamics, with a strong emphasis on electoral campaigns, rallies, and public appearances. It reflects partisan discourse, comparisons between political leaders, reactions to debates, and commentary aligned with conservative media narratives. The content highlights campaign messaging, party unity, and critiques of opposing political figures.

Crime, Security and Justice: focuses on issues related to crime prevention, law enforcement, and the justice system, with particular emphasis on human trafficking, child exploitation, and high-profile criminal investigations. The documents highlight arrests, undercover operations, legal accountability, and public awareness efforts aimed at protecting vulnerable populations and addressing systemic failures.

Foreign Policy: It concentrates on international relations with a strong regional focus on Taiwan and its geopolitical context. It covers elections, governance, cross-strait relations with China, security concerns, international diplomacy, and Taiwan's role in global political institutions, emphasizing political analysis and regional stability.

War: It addresses armed conflict and geopolitical violence, with a primary focus on the Israel–Palestine conflict and broader Middle Eastern tensions. It includes political commentary, historical framing, moral arguments, and discussions of military actions, international responses, and ideological interpretations related to ongoing wars.

To characterize the information environment of YouTube videos shared on Telegram, we examine the distribution of unique videos across macro-topics (see Fig.  3), contrasting political macro-topics (orange) against non-political macro-topics (blue). The distribution is markedly skewed, indicating that shared videos are unevenly distributed across themes. Some macro-topics account for a substantially larger portion of videos, while others remain comparatively limited in scope.

We observe three broad patterns: (1) Niche macro-topics, located at the bottom of Figure 3, exhibit low sharing volume (e.g., Technology, Crime/Security/Justice, Environment, Health, and Hobbies), suggesting more circumscribed thematic presence. (2) Intermediate-volume macro-topics are dominated by politically salient themes (US Politics, War, and Foreign Policy), indicating sustained attention to political content without fully saturating the platform. (3) Mass-audience macro-topics, notably Religion and Entertainment, contain the highest number of distinct shared videos, each surpassing 110k items. These results show that Telegram's political sphere is not exclusively dominated by political content in terms of unique video diversity. However, political macro-topics occupy a prominent position, combining substantial thematic breadth with consistent presence in the corpus of shared links.

Table 3: Mean Perspective API toxicity scores per macro-topic. Asterisk () indicates distributions significantly higher than in the rest of the dataset (Mann–Whitney U test with Bonferroni correction, p < 0.05/60).

Macro-topic

Toxicity

Severe toxicity

Insult

Obscene

Threat

Identity attack

Crime, security and justice

0.100

0.082

0.090

0.088

0.084

0.092

War

0.098

0.084

0.085

0.086

0.088

0.097

U.S. politics

0.068

0.054

0.063

0.057

0.056

0.059

Foreign Policy

0.052

0.043

0.047

0.046

0.045

0.049

Entertainment

0.066

0.056

0.062

0.064

0.056

0.058

Religion

0.055

0.046

0.049

0.049

0.046

0.051

Hobbies

0.046

0.037

0.043

0.041

0.037

0.038

Health

0.042

0.031

0.036

0.033

0.032

0.035

Environment

0.037

0.029

0.032

0.030

0.030

0.031

Technology

0.036

0.028

0.033

0.035

0.028

0.029

4.1.2 Toxicity by Macro-topic. We next quantify the toxicity associated with each macro-topic using the Perspective API. Table 3 reports mean scores across six toxicity-related dimensions. To assess whether toxicity in each macro-topic is statistically higher than in the remainder of the dataset, we apply one-sided Mann–Whitney U tests with Bonferroni correction [34], to compare the distribution within each macro-topic and outside it. The results reveal a consistent pattern: the political macro-topics Crime, Security and Justice, War, and U.S. Politics concentrate the highest mean toxicity levels compared to the rest of the dataset. In particular, Crime, Security and Justice and War display the highest values across all measured dimensions (overall toxicity, severe toxicity, insult, obscene, threat, and identity attack). Substantively, these themes involve sensitive narratives centered on criminal activity, armed conflict, and geopolitical confrontation, which plausibly contribute to elevated toxicity scores. Notably, Crime, Security and Justice is a comparatively low-volume macro-topic, whereas War is substantially more prevalent; yet their mean toxicity levels are comparable. This contrast indicates that toxicity is not proportional to content volume.

Figure 4 provides a distributional view of toxicity. Political macro-topics show broader interquartile ranges, indicating that higher toxicity scores extend across a larger portion of videos. In contrast, non-political macro-topics display more compressed distributions, with high toxicity values concentrated in a smaller subset of items.

To complement the analysis, we performed pairwise comparisons (topic-vs-topic) using the one-sided Mann-Whitney U test with Bonferroni Correction. Figure 5 presents the results matrix, in which dark cells indicate that the toxicity distribution of the macro-topic in the line (L) is significantly higher than that of the macro-topic in the column (C). The results reveal a hierarchy of toxicity across macro-topics. War consistently dominates other categories, followed by Crime, Security and Justice and U.S. Politics, whose distributions are significantly higher than most other themes. More broadly, political macro-topics tend to exhibit higher toxicity distributions than non-political ones. An exception emerges for Entertainment, whose toxicity distribution exceeds that of Foreign Policy, indicating that the political vs non-political distinction does not perfectly align with toxicity levels. The pairwise comparisons indicate that toxicity is systematically stratified across themes rather than uniformly distributed.

Figure 5: Pairwise comparisons between macro-topics using Mann-Whitney U test with Bonferroni Correction ( p < 0.05/90) - Macrotopic L (Line) x Macrotopic (C) (Column)

Answer to RQ1:. Our findings delineate a heterogeneous thematic landscape of YouTube videos shared within U.S. political Telegram chats during 2024. Although these chats are supposed to be politically oriented, the YouTube shared content spans both explicitly political themes and categories not directly related to electoral or governance issues. Political macro-topics exhibit consistently higher toxicity than non-political themes across all measured attributes. Differences between macro-topics persist at the distributional level, indicating that toxicity varies across issue domains rather than being confined to isolated dimensions.

For the remainder of the paper, we focus on political macro-topics and retain three non-political categories as reference groups: Religion, Health, and Hobbies. As the central objective is to analyze political content, we restricted non-political categories to optimize the visualization of subsequent analyses, selecting representatives with different volume and toxicity profiles. Religion is included as a high-volume non-political category with intermediate toxicity, while Health and Hobbies provide a low-toxicity baseline, as they exhibit low mean scores and no significant difference in pairwise testing.

4.2 Telegram Diffusion Dynamics (RQ2)

This section investigates how the macro-topics circulate within Telegram: we analyze how shared YouTube videos spread across chats, compare their reach, reposting frequency, and persistence, and assess whether toxicity is associated with these diffusion patterns. We operationalize diffusion using four complementary metrics computed from URL occurrences: (i) sends, the number of times a given video is posted on Telegram (i.e., the total count of occurrences of its YouTube URL/ID); (ii) unique chats reached, the number of distinct public chats in which the video appears (capturing breadth of exposure across communities); (iii) lifetime, defined as one plus the number of days between the first and the last appearance of the video in the dataset (capturing persistence in the ecosystem) (iv) transfer time, defined as the difference (in number of days) between the date a video is published on YouTube and its first appearance on Telegram. Using these metrics, we compare political vs. non-political macro-topics and examine correlations between diffusion and toxicity attributes.

4.2.1 Information propagation patterns. Figure 6 shows the complementary cumulative distribution of videos as a function of number of sends. The main panel contrasts political vs. non-political content, while the zoomed panel highlights macro-topic differences in the low-send regime (0–10 sends), where most variation concentrates.

Figure 6: Complementary cumulative distribution of videos by number of sends.

Two high-level regularities emerge. First, the majority of videos exhibit limited diffusion. Of the 263,412 videos across the seven retained macro-topics analyzed in this subsection, 164,548 (approximately 62.5%) are posted exactly once. For these items, circulation is minimal: the video appears in a single chat and has a lifetime of one day. This pattern indicates a high churn rate for YouTube content on Telegram and holds for both political and non-political categories. However, meaningful differences surface at the macro-topic level. While U.S. Politics and Crime, Security and Justice have fewer than 50% of their videos posted only once, macro-topics such as Hobbies and Foreign Policy exceed 70% in this condition. This suggests that some themes are more likely to clear the initial “activation” barrier (i.e., to be reposted beyond their first appearance), possibly consistent with topic-dependent engagement and community interest. Given the prevalence of one-off postings, we restrict the remainder of this subsection to videos that are reposted at least once. We define these videos as viral.

Figure 7: Complementary cumulative distributions of diffusion metrics for viral content.

Figure 7 reports complementary cumulative distributions for viral videos across the four diffusion metrics: sends, unique chats, lifetime, and transfer time. For sends and unique chats, political and non-political content exhibit broadly similar distributions, indicating that once the viral threshold is crossed, content can reach comparable breadth and volume of reposts regardless of topic class.

In contrast, lifetime sharply distinguishes the two categories: political videos have substantially shorter circulation windows than non-political videos. Transfer time shows that videos uploaded on YouTube transfer to Telegram in a substantially shorter period of time for political content. The solid grey lines corroborate that this pattern holds across all political macro-topics.

Although political and non-political videos reach comparable breadth once reposted, political content circulates over shorter time spans. In other words, political videos tend to exhibit higher daily intensity, shorter transfer times, and shorter lifetimes—characterized by rapid bursts of diffusion rather than long-tail persistence, and quickly cross platform barriers.

To complement these analyses, Table 4 presents descriptive statistics for ratios of dissemination metrics that help characterize different aspects of propagation dynamics. On average, political videos exhibit higher sends/day and chats/day, indicating that they generate more sends and reach more chats per day than non-political videos. In contrast, non-political content shows a substantially higher standard deviation in sends per day, suggesting a more heterogeneous distribution: most videos have low dissemination, but a few outliers become extremely viral. Political content, by comparison, exhibits a more consistent pattern of virality. In addition, political videos exhibit slightly higher chats/send values, indicating that each video reaches a marginally larger number of distinct chats, suggesting more distributed diffusion.

Table 4: Dissemination metrics of contents

Metric

Politics



Non-Politics




Mean

Std

Median

Mean

Std

Median

Sends / Day

2.484

7.806

2.000

1.685

10.460

1.500

Chats / Day

2.287

2.432

2.000

1.431

1.605

1.000

Chats / Sends

0.937

0.158

1.000

0.884

0.212

1.000

Taken together, these findings indicate that differences between political and non-political content lie in their temporal dynamics. Conditional on virality, political videos tend to reach multiple chats rapidly and diffuse with higher daily intensity over shorter time spans, consistent with burst-like circulation. In contrast, non-political videos exhibit comparatively slower, more sustained diffusion patterns. This might be due to the election period and heightened political interest.

4.2.2 Toxicity–diffusion associations.

Figure 8: Spearman correlations between toxicity attributes (Perspective API) and diffusion metrics, by macro-topic.

We examine the association between toxicity and diffusion across macro-topics using Spearman correlations. Results indicate statistically significant but consistently weak relationships between toxicity attributes and diffusion metrics, with correlation magnitudes generally around 0.1. Figure 8 summarizes the statistically significant associations (p < 0.05). Political macro-topics, particularly Foreign Policy and U.S. Politics, show predominantly positive correlations, suggesting that toxic content tends to circulate slightly more within these domains. By contrast, transfer time follows a somewhat different pattern: correlations are more often positive across macro-topics, indicating that higher-toxicity videos tend to appear on Telegram after longer delays. Overall, these results suggest that toxicity is present across different diffusion patterns, but that its association with diffusion remains limited in magnitude. For other themes, associations are weaker and less systematic, reinforcing that toxicity–diffusion relationships are context-dependent but consistently weak, even when statistically significant.

Answer to RQ2:. We find that diffusion is highly skewed: most videos are posted once, indicating rapid turnover and limited propagation. Among the subset of viral videos (sends > 1), political and non-political content achieve comparable reach and reposting volume. However, political videos display a distinct temporal pattern, with shorter lifetimes and higher daily intensity. In practice, political content diffuses through burst-like, short-duration episodes that quickly reach multiple chats, whereas non-political videos circulate more gradually and persist for longer periods.

Regarding toxicity, correlations with diffusion are consistently small in magnitude. Toxicity is therefore not a dominant determinant of reach or reposting frequency; at most, it behaves as a weak, theme-contingent co-factor (with slightly more positive associations in U.S. Politics and Foreign Policy).

4.3 Temporal Dynamics of Content (RQ3)

Figure 9: Temporal dynamics of activity, participation, and content production across macro-topics in 2024.

Beyond cross-sectional diffusion patterns, we examine how diffusion and toxicity dynamics evolve. We analyze the weekly evolution of diffusion and toxicity throughout 2024, tracking (i) the volume of shared videos, (ii) the number of active chats, and (iii) changes in average toxicity across macro-topics. This temporal perspective allows us to distinguish sustained from episodic patterns of circulation and to assess how sharing dynamics vary in correspondence with high-salience political and informational events.

4.3.1 Weekly volume dynamics. We first count the total number of videos posted per macro-topic on a weekly basis (Figure 9a). Several macro-topics display pronounced spikes throughout 2024, most notably U.S. Politics, as well as Religion, Foreign Policy, and Health. Manual inspection of videos in peak weeks, combined with external news verification, indicates that many of these surges coincide with high-salience events. In U.S. Politics, spikes align with major campaign developments, including the attempted assassination of then-candidate Donald Trump (July 13) and the U.S. presidential election (November 5). Similar event-driven spikes are observable in other macro-topics. In Religion, peaks coincide with the 53rd International Eucharistic Congress and Vatican-related hearings involving Pope Francis. In Foreign Policy, surges align with Benny Gantz's resignation from Israel's war cabinet and the European Parliament elections. Likewise, spikes in Health correspond to the Marburg virus outbreak reported in Rwanda and renewed coverage of mpox.

However, comparing total sends with weekly counts of unique videos (Figure 9b) reveals an important distinction. While peaks in U.S. Politics remain clearly visible under unique-video counting, indicating a broad influx of new content, spikes in Religion, Foreign Policy, and Health become substantially attenuated. This suggests that many of their observed surges are driven primarily by intensive recirculation of a limited set of videos rather than by increased content diversity. In the case of U.S. Politics, by contrast, the persistence of peaks under unique-video counting reinforces the interpretation that, especially during the electoral period, diffusion is sustained by a strong and continuous influx of new content. This pattern is consistent with a faster, shorter-lived attention cycle, in which videos are rapidly introduced, circulated, and displaced as new political developments emerge.

Furthermore, in Figure 9a, U.S. Politics exhibits not only episodic spikes but also a sustained upward shift in baseline volume over the year. During the first half of 2024, weekly counts remain below approximately 3k videos, whereas in the months leading up to the election, they stabilize around 5k. This suggests an intensification in the recirculation of weekly video content on this topic, as Figure 9b does not show a corresponding difference between the two semesters. On the other hand, other macro-topics do not display a comparable increase.

4.3.2 Active chat participation. Figure 9c reports the weekly number of chats in which videos associated with each macro-topic are shared. Most themes exhibit a relatively stable base of participating chats over time, indicating that recurring discussion tends to take place within an established set of chats. Political macro-topics, especially U.S. Politics and Foreign Policy, depart from this pattern during major events: in high-salience weeks, the number of active chats increases sharply, mirroring the volume of spikes observed earlier. These surges suggest that external events temporarily broaden participation beyond the usual set of recurring chats.

Looking at cumulative participation, most macro-topics show a rapid expansion phase early in the year, followed by a marked slowdown in new-chat entry. After the first 7 weeks, already about 50% of the chats that will eventually participate have appeared, which corresponds to roughly 3,300 out of 6,200 chats for the largest macro-topic (Religion) and about 1,100 out of 2,400 for the smallest one (Crime, Security and Justice). From that point onward, expansion becomes substantially incremental, proceeding at only about 1-1.5% of the final total per week, indicating that diffusion is primarily sustained by a relatively stable core of recurring chats.

4.3.3 Weekly toxicity dynamics. We investigate how toxicity levels evolve across macro-topics in 2024. Weekly toxicity is computed as the average overall-toxicity score of videos posted in each week, allowing comparisons independent of posting volume (Figure 9d). Across themes, toxicity remains comparatively stable, with only modest fluctuations. This suggests that spikes in activity are not primarily driven by short-lived concentrations of highly toxic videos. Instead, macro-topics differ in their baseline toxicity levels, while within each macro-topic the average toxicity varies only marginally over time.

This stability is particularly evident for higher-toxicity political macro-topics such as War and Crime, Security and Justice, which remain consistently elevated week after week. Overall, high-salience events appear to modulate how much content circulates and how many chats participate, but they do not substantially alter the average toxicity within each thematic domain.

Answer to RQ3:. Temporal dynamics vary systematically across macro-topics. Three distinct patterns emerge: (i) sustained growth in U.S. Politics before the U.S. election with event-driven amplification; (ii) punctuated, event-linked spikes largely driven by recirculation of a small set of focal videos in Foreign Policy, Health, and Religion; and (iii) relative stability with limited week-to-week variance in Crime/Security/Justice, Hobbies, and War. When restricting the analysis to unique videos, peaks attenuate for event-driven macro-topics, indicating that many bursts reflect repeated sharing of a few items rather than increases in distinct content. In contrast, spikes in U.S. Politics remain visible, consistent with a broader influx of new videos during major events.

Active-chat patterns further show that most themes rely on a stable base of participating chats, while high-visibility political events temporarily broaden participation. Finally, weekly toxicity remains largely stable within each macro-topic: events amplify volume and participation, but do not substantially alter average toxicity levels.

5 Conclusion and Future Work

This study conceptualized Telegram as a cross-platform redistribution layer in which shared URLs function as observable traces of how external media content is recontextualized and amplified. By analyzing YouTube videos shared in Telegram chats throughout 2024, we characterized the thematic structure of circulating content, quantified toxicity, and examined diffusion in terms of reach, recirculation, persistence, and temporal transfer.

Three main findings emerge. First, the thematic landscape is heterogeneous: although political content is prominent, politically oriented chats also circulate substantial non-political material. Moreover, political macro-topics consistently exhibit higher toxicity than non-political ones. Second, diffusion is highly skewed. Most videos are shared once and disappear quickly. Among reposted videos, political and non-political content achieve comparable reach, but political videos spread over compressed time horizons, with shorter lifetimes and higher daily intensity. The key difference lies not in the ultimate scale of diffusion, but in its temporal structure. Third, toxicity exhibits small but statistically significant associations with diffusion metrics in selected macro-topics. These effects are not systematic, however, and do not uniformly translate into greater reach across themes. At the same time, toxicity levels remain comparatively stable over time within each macro-topic, suggesting that toxicity is primarily structured by thematic domain rather than driven by episodic spikes in posting volume. Taken together, these results suggest that cross-platform amplification is shaped more by topic and temporality than by toxicity alone. Shared links do not simply transport content across platforms; as cross-platform bridging acts, they reorganize its visibility along thematic and temporal lines, producing burst-like amplification in some domains and stable recirculation in others.

Future work can extend this perspective by modeling URL sharing as explicit cascades across group-level interaction networks, incorporating richer video signals beyond metadata, and employing causal designs to disentangle topic salience from content attributes. A deeper examination of feedback loops between Telegram bursts and subsequent YouTube engagement would further clarify how cross-platform linking contributes to the broader dynamics of online political communication.

6 Limitations and Ethical Considerations

We acknowledge several limitations of this study. First, toxicity is measured using video metadata (title and description) rather than full transcripts or audiovisual content. Although this enables large-scale analysis, it may not capture contextual nuance or implicit forms of harmful speech present in the videos themselves.

Second, our analysis relies on an existing large-scale dataset of Telegram chats related to the 2024 U.S. election. While the dataset is extensive in scope, it was not collected by the authors, and we therefore do not claim comprehensive coverage of Telegram activity. In addition, because the corpus was originally constructed using politically oriented seed terms in the context of the election, it may introduce sampling biases in both topical coverage and temporal dynamics. At the same time, prior work has shown that the snowball-based collection process naturally expands beyond strictly political chats. Therefore, our results should be interpreted in light of this potential bias: they characterize patterns within the observed corpus rather than the entire platform, while still enabling meaningful comparisons between political and non-political content.

Third, the relationship between toxicity and diffusion is assessed using correlational methods. These findings do not imply causal effects of toxicity on amplification dynamics.

From an ethical perspective, the study analyzes publicly available content and focuses exclusively on aggregate patterns. No attempt was made to identify individual users or profile-specific groups. Toxicity scores derived from the Perspective API are probabilistic indicators of linguistic features and should not be interpreted as normative judgments about content or intent. Care was taken to avoid reproducing harmful material beyond what is necessary for analysis.

Acknowledgments

This work was supported by CNPq, CAPES, INCT-TILDIAR (grant no. 408490/2024-1), FAPEMIG, AWS, FAPESP and the Spoke 1 “FutureHPC & BigData" of ICSC — Centro Nazionale di Ricerca in High-Performance-Computing, Big Data and Quantum Computing, funded by the European Union — NextGenerationEU.

Notes

1From this point onward, the term ‘chat’ will be used to refer to both groups (many-to-many communication) and channels (one-to-many).


3Even limited to 128 tokens [44], it presents better results than models with more tokens

Source


    Imported from ACM’s structured HTML source. ACM Reference Format: Mateus Brito, Ester Souza, Thiago Braga, Giordano Paoletti, Luca Vassio, Jussara M. Almeida, and Leonardo Rocha. 2026. Link-Traced Amplification: How Telegram Redistributes YouTube Political Content at Scale. In 37th ACM Conference on Hypertext (HT '26), September 14--18, 2026, London, United Kingdom. ACM, New York, NY, USA 12 Pages. https://doi.org/10.1145/3800935.3830834

References

[1] Lorenzo Alvisi, Serena Tardelli, and Maurizio Tesconi. 2025. Mapping the Italian Telegram Ecosystem: Communities, Toxicity, and Hate Speech. arXiv e-prints (2025), arXiv–2504.

[2] Katarina Bader, Kathrin Friederike Müller, and Lars Rinsdorf. 2025. Wanderer between the worlds: Telegram use from the users’ perspective. Publizistik 70, 1 (2025), 133–155. https://doi.org/10.1007/s11616-025-00874-x

[3] Steven Bird, Ewan Klein, and Edward Loper. 2009. Natural language processing with Python: analyzing text with the natural language toolkit. " O'Reilly Media, Inc.".

[4] Lennart Björneborn and Peter Ingwersen. 2004. Toward a basic framework for webometrics. Journal of the American society for information science and technology 55, 14 (2004), 1216–1227.

[5] Leonardo Blas, Luca Luceri, and Emilio Ferrara. 2025. Unearthing a Billion Telegram posts about the 2024 US Presidential Election: Development of a public dataset. In Proc. Comp. of The Web Conf ’25.

[6] Alexandre Bovet and Peter Grindrod. 2022. Organization and evolution of the UK far-right network on Telegram. Applied Network Science (2022).

[7] Leo Breiman. 2001. Random Forests. Mach. Learn. 45, 1 (Oct. 2001), 5–32. https://doi.org/10.1023/A:1010933404324

[8] Victor S. Bursztyn and Larry Birnbaum. 2019. Thousands of Small, Constant Rallies: A Large-Scale Analysis of Partisan WhatsApp Groups. In 2019 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM). 484–488. https://doi.org/10.1145/3341161.3342905

[9] Andrew Chadwick. 2017. The hybrid media system: Politics and power. Oxford University Press.

[10] Johnny Chan and Yuming Li. 2024. Unveiling disguised toxicity: A novel pre-processing module for enhanced content moderation. MethodsX 12 (June 2024), 102668. https://doi.org/10.1016/j.mex.2024.102668

[11] Matthew Childs, Cody Buntain, Milo Z. Trujillo, and Benjamin D. Horne. 2022. Characterizing youtube and bitchute content and mobilizers during us election fraud discussions on twitter. In Proceedings of the 14th ACM Web Science Conference 2022. 250–259.

[12] Federico Cinus, Marco Minici, Luca Luceri, and Emilio Ferrara. 2025. Exposing cross-platform coordinated inauthentic activity in the run-up to the 2024 us election. In Proc.The Web Conf’25.

[13] Adji B Dieng, Francisco JR Ruiz, and David M Blei. 2020. Topic modeling in embedding spaces. Transactions of the Association for Computational Linguistics 8 (2020).

[14] Carlos Henrique Gomes Ferreira, Fabricio Murai, Ana P. C. Silva, Martino Trevisan, Luca Vassio, Idilio Drago, Marco Mellia, and Jussara M. Almeida. 2022. On network backbone extraction for modeling online collective behavior. PLoS ONE 17(9) (2022).

[15] Paula Fortuna and Sérgio Nunes. 2018. A survey on automatic detection of hate speech in text. Acm Computing Surveys (Csur) 51, 4 (2018), 1–30.

[16] Valerio La Gatta, Luca Luceri, Francesco Fabbri, and Emilio Ferrara. 2023. The interconnected nature of online harm and moderation: Investigating the cross-platform spread of harmful content between youtube and twitter. In Proceedings of the 34th ACM conference on hypertext and social media. 1–10.

[17] Bryan T. Gervais. 2025. Incivility or Invalidity? Evaluating Perspective API Scores as a Measure of Political Incivility. The International Journal of Press/Politics (2025). https://doi.org/10.1177/1532673X241309627

[18] Lillie Godinez and Eni Mustafaraj. 2024. Youtube and conspiracy theories: A longitudinal audit of information panels. In Proceedings of the 35th ACM Conference on Hypertext and Social Media. 273–284.

[19] Google. 2026. YouTube Data API v3. Retrieved March 18, 2026. https://developers.google.com/youtube/v3

[20] Maarten Grootendorst. 2022. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv:2203.05794 (2022).

[21] Scott A Hale, Adriano Belisario, Ahmed Nasser Mostafa, and Chico Camargo. 2024. Analyzing Misinformation Claims During the 2022 Brazilian General Election on WhatsApp, Twitter, and Kwai. International Journal of Public Opinion Research 36, 3 (07 2024), edae032. arXiv:https://academic.oup.com/ijpor/article-pdf/36/3/edae032/58513480/edae032.pdf https://doi.org/10.1093/ijpor/edae032

[22] Hans WA Hanley and Zakir Durumeric. 2024. Partial mobilization: Tracking multilingual information flows amongst Russian media outlets and Telegram. In Proc. ICWSM’24.

[23] Laura Hanu and Unitary team. 2020. Detoxify. Github. https://github.com/unitaryai/detoxify.

[24] Aliaksandr Herasimenka, Jonathan Bright, Aleksi Knuutila, and Philip N Howard. 2023. Misinformation and professional news on largely unmoderated platforms: The case of Telegram. Journal of Information Technology & Politics 20, 2 (2023), 198–212.

[25] Lars Heyn. 2025. Drilling a Telegram–YouTube Pipeline: Cross-Platform Affordances of German Anti-Government Extremist Networks. Studies in Conflict & Terrorism (2025). https://doi.org/10.1080/1057610X.2025.2595842

[26] Matthew Honnibal, Ines Montani, Sofie Van Landeghem, Adriane Boyd, et al. 2020. spaCy: Industrial-strength natural language processing in python. (2020).

[27] Mohamad Hoseini, Philipe de Freitas Melo, Fabrício Benevenuto, Anja Feldmann, and Savvas Zannettou. 2024. Characterizing Information Propagation in Fringe Communities on Telegram. In International AAAI Conference on Web and Social Media.

[28] Mohamad Hoseini, Philipe Melo, Fabricio Benevenuto, Anja Feldmann, and Savvas Zannettou. 2023. On the Globalization of the QAnon Conspiracy Theory Through Telegram. In Proc.ACM WebSci’23.

[29] Alexander Hoyle, Pranav Goel, Denis Peskov, Andrew Hian-Cheong, Jordan Boyd-Graber, and Philip Resnik. 2021. Is automated topic model evaluation broken? The incoherence of coherence. In Proc. NeurIPS’21. Article 155.

[30] Massimo La Morgia, Alessandro Mei, and Alberto Maria Mongardini. 2025. TGDataset: Collecting and Exploring the Largest Telegram Channels Dataset. In Proc. ACM SIGKDD’25.

[31] Alyssa Lees, Vinh Q. Tran, Yi Tay, Jeffrey Sorensen, Jai Gupta, Donald Metzler, and Lucy Vasserman. 2022. A New Generation of Perspective API: Efficient Multilingual Character-level Transformers. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Washington DC, USA) (KDD ’22). Association for Computing Machinery, New York, NY, USA, 3197–3207. https://doi.org/10.1145/3534678.3539147

[32] Yanli Liu, Yourong Wang, and Jian Zhang. 2012. New machine learning algorithm: Random forest. In International conference on information computing and applications. Springer, 246–252.

[33] Marcelo Sartori Locatelli, Josemar Caetano, Wagner Meira Jr, and Virgilio Almeida. 2022. Characterizing vaccination movements on YouTube in the United States and Brazil. In Proceedings of the 33rd ACM Conference on Hypertext and Social Media. 80–90.

[34] H. B. Mann and D. R. Whitney. 1947. On a test of whether one of two random variables is stochastically larger than the other. The Annals of Mathematical Statistics 18, 1 (1947), 50–60. https://doi.org/10.1214/aoms/1177730491

[35] Breno Matos, Rennan C Lima, Jussara M Almeida, Marcos A Gonçalves, and Rodrygo LT Santos. 2024. Intellectual dark web, alt-lite and alt-right: Are they really that different? a multi-perspective analysis of the textual content produced by contrarians. Social Network Analysis and Mining 14, 1 (2024), 32.

[36] David Mimno, Hanna Wallach, Edmund Talley, Miriam Leenders, and Andrew McCallum. 2011. Optimizing semantic coherence in topic models. In Proc. EMNLP ’11.

[37] Shuyo Nakatani. 2010. langdetect. Accessed: 2025-05-12. https://pypi.org/project/langdetect/

[38] OpenAI. 2026. Introducing GPT-5.4. https://openai.com/index/introducing-gpt-5-4/. GPT-5.4 used in ChatGPT as GPT-5.4 Thinking. Accessed: 2026-03-12.

[39] Giordano Paoletti, Jussara M. Almeida, Luca Vassio, Marcos André Gonçalves, and Marco Mellia. 2025. Join the Chat: How Curiosity Sparks Participation in Telegram Groups. In Proc. ICWSM’25.

[40] Giordano Paoletti, Carlos HG Ferreira, Luca Vassio, Leonardo Rocha, and Jussara M Almeida. 2025. Tracing the 2024 US election debate on Telegram with LLMs and graph analysis. Social Network Analysis and Mining 15, 1 (2025), 91.

[41] Aakash Parmar, Rakesh Katariya, and Vatsal Patel. 2018. A review on random forest: An ensemble classifier. In International conference on intelligent data communication technologies and internet of things. Springer, 758–763.

[42] John Pavlopoulos, Leo Laugier, Alexandros Xenos, Jeffrey Sorensen, and Ion Androutsopoulos. 2022. From the Detection of Toxic Spans in Online Discussions to the Analysis of Toxic-to-Civil Transfer. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dublin, Ireland, 3721–3734. https://doi.org/10.18653/v1/2022.acl-long.259

[43] Alessandro Perlo, Giordano Paoletti, Nikhil Jha, Luca Vassio, Jussara Almeida, and Marco Mellia. 2025. Topic-wise Exploration of the Telegram Group-verse. In Proc. Comp. ACM The Web Conf.’25.

[44] Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics. http://arxiv.org/abs/1908.10084

[45] Gustavo Resende, Philipe Melo, Hugo Sousa, Johnnatan Messias, Marisa Vasconcelos, Jussara Almeida, and Fabrício Benevenuto. 2019. (Mis) information dissemination in WhatsApp: Gathering, analyzing and countermeasures. In The World Wide Web Conference.

[46] Manoel Horta Ribeiro, Raphael Ottoni, Robert West, Virgílio A. F. Almeida, and Wagner Meira. 2020. Auditing radicalization pathways on YouTube. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (Barcelona, Spain). 131–141.

[47] Peter J Rousseeuw. 1987. Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. J. Comput. Appl. Math. 20 (1987).

[48] Karla Schäfer and Jeong-Eun Choi. 2023. Transparency in Messengers: A Metadata Analysis Based on the Example of Telegram. In Proceedings of the 34th ACM Conference on Hypertext and Social Media (Rome, Italy) (HT ’23). Association for Computing Machinery, New York, NY, USA, Article 12, 3 pages. https://doi.org/10.1145/3603163.3609034

[49] Heidi Schulze, Julian Hohner, Simon Greipl, Maximilian Girgnhuber, Isabell Desta, and Diana Rieger. 2022. Far-right conspiracy groups on fringe platforms: a longitudinal analysis of radicalization dynamics on Telegram. Convergence: The International Journal of Research into New Media Technologies 28, 4 (2022).

[50] Elin Thorgren, Alireza Mohammadinodooshan, and Niklas Carlsson. 2024. Temporal Dynamics of User Engagement on Instagram. In Proc. ACM WebSci ’24.

[51] Aleksandra Urman and Stefan Katz. 2022. What they do in the shadows: examining the far-right networks on Telegram. Inf. Comm. & Soc. 25, 7 (2022), 904–923.

[52] Otavio R Venâncio, Carlos HG Ferreira, Jussara M Almeida, and Ana Paula C da Silva. 2024. Unraveling user coordination on Telegram: A comprehensive analysis of political mobilization during the 2022 Brazilian Presidential election. In Proc. ICWSM’24.

[53] Samantha Walther and Andrew McCoy. 2021. US extremism on Telegram. Perspectives on Terrorism 15, 2 (2021), 100–124.

[54] Yunkang Yang, Ramesh Paudel, Jordan McShan, Matthew Hindman, H Howie Huang, and David Broniatowski. 2025. Coordinated link sharing on Facebook. SciRep 15, 1 (2025), 15684.

[55] Savvas Zannettou, Tristan Caulfield, Jeremy Blackburn, Emiliano De Cristofaro, Michael Sirivianos, Gianluca Stringhini, and Guillermo Suarez-Tangil. 2018. On the origins of memes by means of fringe web communities. In Proceedings of the internet measurement conference 2018. 188–202.

[56] Savvas Zannettou, Tristan Caulfield, Emiliano De Cristofaro, Nicolas Kourtelris, Ilias Leontiadis, Michael Sirivianos, Gianluca Stringhini, and Jeremy Blackburn. 2017. The web centipede: understanding how web communities influence each other through the lens of mainstream and alternative news sources. In Proceedings of the 2017 internet measurement conference. 405–417.

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime