Transparency in Messengers - A Metadata Analysis Based on the Example of TelegramA short paper deriving privacy-friendly metadata key figures from 770 German disinformation-related Telegram channels and groups, finding that channels drive information spread while group-originated posts are never forwarded.

Transparency in Messengers - A Metadata Analysis Based on the Example of Telegram

Published in HT '23: 34th ACM Conference on Hypertext and Social Media · DOI: 10.1145/3603163.3609034 · License: © Copyright held by the owner/author(s).

Authors: Karla Schäfer, Jeong-Eun Choi

Abstract

Social media platforms and messenger services such as Telegram have created a new way of publishing and consuming news. To combat the resulting negative effects, such as disinformation, social media analytics (SMA) can be applied to, inter alia, reconstruct the spread of information across platforms or to identify key actors and interactions between users. This paper examines metadata in Telegram to provide a privacy-friendly basis for further SMA. The characteristics of 770 Telegram channels and groups were derived by extracting initial key figures. We observed that messages that were created in channels were forwarded to other actors, while not a single original message created in groups was forwarded. This suggests that channels have a much greater impact on generating and spreading information to other Telegram actors than groups.

Ccs Concepts

• Information systems →Web and social media search; • Security and privacy →Social network security and privacy.

Keywords

Telegram, Metadata, Social Media Analytics, Privacy

ACM Reference Format: Karla Schäfer and Jeong-Eun Choi. 2023. Transparency in Messengers - A Metadata Analysis Based on the Example of Telegram. In 34th ACM Conference on Hypertext and Social Media (HT ’23), September 4–8, 2023, Rome, Italy. ACM, New York, NY, USA, 3 pages. https://doi.org/10.1145/ 3603163.3609034

1 Introduction

The shift of information sources from traditional media to various social media platforms facilitated the spread of information to an unprecedented scale that consequently elevated the danger of possible misuse of digital communication[5, 6, 8, 14]. Using social media analytics (SMA)[19], some of these issues can be mitigated. For instance, SMA allows identifying echo chambers[4, 21] or spreaders of disinformation[10, 18, 22, 23]. However, applying SMA to real-world use cases presents some challenges, often caused by

the volume, variety, veracity, and velocity of the data[2, 9, 20]. Telegram[11, 12] is a messenger application that is growing in popularity, e.g. among users who have been banned or who moved away from leading social media outlets like Facebook, Instagram and Twitter[1, 15] due to comparatively fewer regulations. Telegram consists of private and public groups, channels and individual chats, and became known, among other things, for offering a high level of anonymity to its users1. This paper examines metadata accessed through the Telegram API to obtain key figures that allow to accurately retrieve insights of channels and groups, which can be further used to identify information dissemination patterns to combat disinformation or create transparency for users to counteract the formation of echo chambers[3, 7, 13, 17]. Analysing the metadata can provide an opportunity for non-invasive data analysis, as solely data about the group/channel characteristics are extracted and not the content.

2 Dataset

A dataset consisting of 770 Telegram channels and groups (511 channels, 259 groups) crawled via the Telegram API in the period from 25-03-2022 to 28-02-2023 was created. These 770 public channels and groups were identified in 15 interviews with journalism and fact-checking experts to identify key actors in the spread of disinformation in Germany. The interviews were conducted as part of the project DYNAMO2 and organized by our partner "Hochschule der Medien Stuttgart" [16]. Due to applicable privacy policies, the exact names and IDs of the channels and groups cannot be disclosed. Thematically, the channels/groups considered include conspiracy theories, radical right-wing ideas and vaccination opponents. While we examined all identified channels and groups, the analysis on two channels (𝐶ℎ1,𝐶ℎ2) and one group (𝐺𝑟1) are presented as examples.

3 Key Figures

The key figures derived here were extracted using statistical methods. Each posted message is linked to an ID that is unique to its group/channel (counted: totalposts). In some cases, messages are deleted after they have been posted; if this occurs, an ID is missing in the numerical sequence (counted: deletedposts). We suspect that a high number of deleted posts compared to existing posts could serve as a first indicator of suspicious channels/groups, as presumably illegal/compromising content is more likely to be deleted after posting. The number of members/subscribers in groups/channels

1https://telegram.org/blog/ultimate-privacy-topics-2-0 2https://www.forschung-it-sicherheit-kommunikationssysteme.de/projekte/dynamo

Table 1: Metadata analysis results (bold: features of particular interest).

Example Channel/Group Channel𝑎𝑙𝑙 Group𝑎𝑙𝑙 Key Figure Ch1 Ch2 Gr1 min mean max min mean max

totalposts (count posts) 3,454 59,371 1,085 10 5,261 141,724 6 16,122 467,768 deletedposts (count posts) 141 1,087 50,979 0 264 13,126 0 11,630 1,314,775 countparticipants 254,282 197,017 15,601 28 21,492 254,282 3 1,126 25,023 postsperdayavg (avg count posts) 7.49 108.94 1.22 0.02 11.82 242.26 0.01 43.74 1,496.8 realpostsperdayavg (avg count posts) 7.8 110.93 58.37 0.07 12.28 242.8 0.64 71.68 5,092.98 totaltimeonline (days) 461 545 892 289 711 2,610 292 575 1,607

activityposts (%) 0.2% 0.46% 48.39% 0% 1.2% 28.3% 0% 9.3% 98.54% originalposts (%) 99.77% 74.74% 37.33% 0.04% 63.4% 100% 0.49% 50.02% 99% forwardedposts (%) 0.03% 24.8% 14.29% 0% 35.45% 99.9% 0.04% 40.67% 99.26%

originalpostswithforwards (%) 99.68% 92.33% 0% 1.57% 82.68% 100% 0% 0% 0% avgforwardsperoriginalpost (avg count posts) 662.35 252.63 0 0.1 125.2 3,574.09 0 0 0 anonymforwardedposts (%) 0% 0.52% 11.61% 0% 5.54% 74.92% 0% 4.54% 30.65% selfforward (count posts) 0 11,267 0 0 98 15,772 0 1 166

can also be determined (countparticipants). The average number of posts per day is also stated in the metadata (postsperdayavg) and can provide initial information, e.g. to identify inactive or rarely frequented groups/channels, whereby it should be noted that deleted posts are not included in the calculation (deleted posts included: realpostsperdayavg). The total number of days from when the group/channel was created to when the crawling concluded can also be viewed (totaltimeonline), showing the lifespan. In Telegram, a distinction can be made between content posts, which contain the user’s messages, and activity posts (activityposts). Activity posts describe activities within the group/channel such as adding new members or editing the channel title. The content posts can be divided into original posts, i.e. posts created in this group/channel (originalposts), and forwarded posts, i.e. posts created by other actors (users, groups or channels) and forwarded to this group/channel (forwardedposts). Whether the original posts were forwarded at least once to other actors (originalpostswithforwards) can be identified using the "forwards" metadata. This forwards metadata stores the number of users who have forwarded the message; this counter is not incremented when a message is forwarded to itself (selfforward). The metric avgforwardsperoriginalpost reflects this number of forwards as the averaged sum of all forwards of original messages. These two metrics can be used, for example, to determine the reach or even the influence of the content posted by that group/channel. When messages are forwarded, the origin of the message, i.e. the actor where the message was first posted ("userid" or "channelid"), is stored. However, users can disable this in their account settings so that no origin ID is stored for the message (anonymforwardedposts). If the origin ID of the forwarded message is known, the origin can be identified, and a connection can be drawn between these two actors (e.g. as part of a network analysis). The exact path between three actors, i.e. when someone forwards an already forwarded message, cannot be deduced in Telegram. The same ID can also be used to detect whether messages originated from the group/channel have been forwarded to itself again (repost). In this case, the "channelid" of the forwarded message is the same as that of the group/channel currently being viewed (selfforward), indicating possible content monitoring by the administrator or other users who may be trying to draw attention to certain messages by reposting them repeatedly.

4 Discussion

The results (Table 1) show that none of the groups featured original posts with forwards, i.e., none of the posts created in the groups were forwarded to other channels/groups (originalpostswith forwards). In contrast, the channels contain on average 82.68% of original messages with forwards, showing that channels used primarily by users consuming information, have a much greater impact on other actors than groups. The percentage of anonymised forwarded posts is rather low, which allows the origin of forwarded messages to be identified. Another interesting finding poses the number of deleted posts; within the whole dataset this is much higher in the groups than in the channels. This is also visible in 𝐺𝑟1, with 50,979 deleted posts and 1,085 available posts. 𝐶ℎ1 contains an aboveaverage amount of original posts (99.77%), 99.68% of original posts with forwards and 662.35 average forwards per original post. Therefore, this group has a high influence on other Telegram actors. 𝐶ℎ2 features by far the most posts on average per day with 110.93/108.94 posts published. Since only administrators can publish messages in channels, this suggests that 𝐶ℎ2 is run by multiple administrators. 𝐶ℎ2 consist of 24.8% forwarded messages (14,725 posts) of which 76.5% (11,267 posts) are reposts (selfforward). The high number of postsperdayavg and selfforward are strong indicator that 𝐶ℎ2 is actively applying the reposting strategy to foster the dissemination of the original posts.

5 Conclusion

Identifying unusual behaviour through metadata allows selecting channels and groups that need further investigation, which is an important factor in increasing the efficiency and timeliness of SMA. Our goal was to extract information from the metadata and process it in such a way that it could be used as additional context for SMA without using user specific data. We hypothesized in this work that a high number of deleted posts compared to existing posts could serve as a first indicator of suspicious channels/groups. In future work, we will further investigate this assumption.

Acknowledgments

This research work was supported by the National Research Center for Applied Cybersecurity ATHENE as well as within the German Federal Ministry of Education and Research project DYNAMO.

References

[1] Saharsh Agarwal, Uttara M Ananthakrishnan, and Catherine E Tucker. 2022. Deplatforming and the Control of Misinformation: Evidence from Parler. Available at SSRN (2022).

[2] Gema Bello-Orgaz, Jason J Jung, and David Camacho. 2016. Social big data: Recent achievements and new challenges. Information Fusion 28 (2016), 45–59.

[3] Devipsita Bhattacharya and Sudha Ram. 2012. Sharing news articles using 140 characters: A diffusion analysis on Twitter. In 2012 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining. IEEE, 966–971.

[4] Matteo Cinelli, Gianmarco De Francisci Morales, Alessandro Galeazzi, Walter Quattrociocchi, and Michele Starnini. 2021. The echo chamber effect on social media. Proceedings of the National Academy of Sciences 118, 9 (2021), e2023301118. https://doi.org/10.1073/pnas.2023301118 arXiv:https://www.pnas.org/doi/pdf/10.1073/pnas.2023301118

[5] Yuriy Gorodnichenko, Tho Pham, and Oleksandr Talavera. 2021. Social media, sentiment and public opinions: Evidence from# Brexit and# USElection. European Economic Review 136 (2021), 103772.

[6] Andreas Jungherr and Ralph Schroeder. 2021. Disinformation and the Structural Transformations of the Public Arena: Addressing the Actual Challenges to Democracy. Social Media + Society 7, 1 (2021), 2056305121988928. https: //doi.org/10.1177/2056305121988928

[7] Haewoon Kwak, Changhyun Lee, Hosung Park, and Sue Moon. 2010. What is Twitter, a Social Network or a News Media?. In Proceedings of the 19th International Conference on World Wide Web (Raleigh, North Carolina, USA) (WWW ’10). Association for Computing Machinery, New York, NY, USA, 591–600. https://doi.org/10.1145/1772690.1772751

[8] David M. J. Lazer, Matthew A. Baum, Yochai Benkler, Adam J. Berinsky, Kelly M. Greenhill, Filippo Menczer, Miriam J. Metzger, Brendan Nyhan, Gordon Pennycook, David Rothschild, Michael Schudson, Steven A. Sloman, Cass R. Sunstein, Emily A. Thorson, Duncan J. Watts, and Jonathan L. Zittrain. 2018. The science of fake news. Science 359, 6380 (2018), 1094–1096. https://doi.org/10.1126/science. aao2998

[9] Andrew McAfee, Erik Brynjolfsson, Thomas H Davenport, DJ Patil, and Dominic Barton. 2012. Big data: the management revolution. Harvard business review 90, 10 (2012), 60–68.

[10] Erxue Min, Yu Rong, Yatao Bian, Tingyang Xu, Peilin Zhao, Junzhou Huang, and Sophia Ananiadou. 2022. Divide-and-conquer: Post-user interaction network for fake news detection on social media. In Proceedings of the ACM Web Conference 2022. 1148–1158.

[11] Mehrnoosh Mirtaheri, Sami Abu-El-Haija, Fred Morstatter, Greg Ver Steeg, and Aram Galstyan. 2019. Identifying and Analyzing Cryptocurrency Manipulations in Social Media. CoRR abs/1902.03110 (2019). arXiv:1902.03110 http://arxiv.org/ abs/1902.03110

[12] Massimo La Morgia, Alessandro Mei, Alberto Maria Mongardini, and Jie Wu. 2021. Uncovering the Dark Side of Telegram: Fakes, Clones, Scams, and Conspiracy Movements. CoRR abs/2111.13530 (2021). arXiv:2111.13530 https://arxiv.org/ abs/2111.13530

[13] Arash Dargahi Nobari, Malikeh Haj Khan Mirzaye Sarraf, Mahmood Neshati, and Farnaz Erfanian Daneshvar. 2021. Characteristics of viral messages on Telegram; The world’s largest hybrid public and private messenger. Expert systems with applications 168 (2021), 114303.

[14] Yasmim Mendes Rocha, Gabriel Acácio de Moura, Gabriel Alves Desidério, Carlos Henrique de Oliveira, Francisco Dantas Lourenço, and Larissa Deadame de Figueiredo Nicolete. 2021. The impact of fake news on social media and its influence on health during the COVID-19 pandemic: A systematic review. Journal of Public Health (2021), 1–10.

[15] Richard Rogers. 2020. Deplatforming: Following extreme Internet celebrities to Telegram and alternative social media. European Journal of Communication 35, 3 (2020), 213–229.

[16] Tahireh Setz, Inna Vogel, Martin Steinebach, Katarina Bader, nicole krämer, Gerrit Hornung, York Yannikos, Lars Rinsdorf, Jan Kluck, and Carolin Jansen. 2022. Desinformationen und Messengerdienste: Herausforderung und Lösungsansätze. 317–370. https://doi.org/10.5771/9783748913344-317

[17] Kai Shu, Suhang Wang, and Huan Liu. 2018. Understanding user profiles on social media for fake news detection. In 2018 IEEE conference on multimedia information processing and retrieval (MIPR). IEEE, 430–435.

[18] Kai Shu, Suhang Wang, and Huan Liu. 2019. Beyond news contents: The role of social context for fake news detection. In Proceedings of the twelfth ACM international conference on web search and data mining. 312–320.

[19] Stefan Stieglitz, Linh Dang-Xuan, Axel Bruns, and Christoph Neuberger. 2014. Social media analytics: An interdisciplinary approach and its implications for information systems. Business & Information Systems Engineering 6 (2014), 89–96.

[20] Stefan Stieglitz, Milad Mirbabaie, Björn Ross, and Christoph Neuberger. 2018. Social media analytics – Challenges in topic discovery, data collection, and data preparation. International Journal of Information Management 39 (2018), 156–168. https://doi.org/10.1016/j.ijinfomgt.2017.12.002

[21] Giacomo Villa, Gabriella Pasi, and Marco Viviani. 2021. Echo chamber detection and analysis: a topology-and content-based approach in the COVID-19 scenario. Social Network Analysis and Mining 11, 1 (2021), 78.

[22] Inna Vogel and Meghana Meghana. 2020. Detecting fake news spreaders on twitter from a multilingual perspective. In 2020 IEEE 7th International Conference on Data Science and Advanced Analytics (DSAA). IEEE, 599–606.

[23] Zilong Zhao, Jichang Zhao, Yukie Sano, Orr Levy, Hideki Takayasu, Misako Takayasu, Daqing Li, Junjie Wu, and Shlomo Havlin. 2020. Fake news propagates differently from real news even at early stages of spreading. EPJ data science 9, 1 (2020), 7.

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime