Can LLMs Beat Humans on Discerning Human-written and LLM-generated Science News?The paper introduces SANews, a manually annotated dataset of paired human-written and GPT-3.5-generated science news articles from ScienceAlert, and shows that a Guided Few-shot prompting template with a single example boosts LLMs (including open-weight LLaMA-3 70b) to match or exceed graduate-stude

Can LLMs Beat Humans on Discerning Human-written and LLM-generated Science News?

Dominik Soós (Computer Science, Old Dominion University, Norfolk, VA, USA, dsoos001@odu.edu), Meng Jiang (Computer Science and Engineering, University of Notre Dame, Notre Dame, IN, USA, mjiang2@nd.edu), and Jian Wu (Computer Science, Old Dominion University, Norfolk, VA, USA, fanchyna@gmail.com)

Published in the 36th ACM Conference on Hypertext and Social Media (HT '25) · DOI: 10.1145/3720553.3746674 · License: CC BY 4.0

Abstract

Science news is increasingly important in connecting scientists and the public by sharing discoveries and innovations. With the rise of large language models (LLMs), there is potential to automate science news creation, but concerns exist about the quality of LLM-generated news versus human-written news. This paper explores whether LLMs can outperform humans in distinguishing between human-written and LLM-generated news. Inspired by the Chain-of-Thought prompting method, we designed a simple yet effective variant called Guided Few-shot (GFS), which encodes the characteristics of news of two types with examples. Our experiments indicated that GFS with just a single example effectively boosted the performance of all LLMs and made open-weight LLMs achieve or exceed the performance of humans and commercial LLMs. The code and data are available in our repository at https://github.com/lamps-lab/sanews.

CCS Concepts


    Computing methodologiesParallel computing methodologies; Natural language generation; Natural language processing.

Keywords

science news detection, machine generated text evaluation, large language model

1 Introduction

Nowadays increasing numbers of people engage with scientific and technological advancements through news outlets and social media. According to a study by Pew Research Center, overall, about 36% of Americans get science news at least a few times a week [12]. Many science news outlets, such as ScienceAlert [36] hire professional editors to write science news articles based on research papers published in peer-reviewed journals. On the other hand, Large Language Models (LLMs), such as GPT series [1, 5, 20], Claude 3 family [3], LLaMA series [14], and have opened new opportunities for generating and evaluating comprehensive textual content across domains [26, 29, 38], including news [10]. However, LLM-generated news has been found to contain hallucinated and biased content [11, 19], raising concerns about its reliability compared to human-written content.

Figure 1: Workflow for human and LLM evaluators to discern human-written and LLM-generated news based on the same scientific paper.

There has been extensive research on AI-generated text detection using black-box and white-box approaches [39]. An experiment compared LLM-generated and human-written social media ads to uncover systematic differences that audiences readily perceive [2]. Work on LLM-assisted educational hypercomics shows that machine-generated stories retain distinctive stylistic fingerprints, making them suitable for authorship-origin studies [15]. Different from social media ads or stories, science news poses a special challenge because the content may include domain knowledge. Several tools [32, 43] have been developed attempting to distinguish LLM-generated text content from human-written text by leveraging differences between their properties. How well general AI text detectors work on science news and whether LLM-based methods can outperform these text detectors and human's performance are open questions. In this paper, we would like to address these open questions. Our contributions are summarized below.


    We find that it is possible for open-source LLMs to achieve or exceed the performance of graduate students and state-of-the-art AI text detectors on the task of discerning human-written and LLM-generated science news.

    We developed a new dataset called SANews consisting of 362 science news articles collected from ScienceAlert as triplets of human-written news, paper abstract, LLM-generated news. The dataset provides a first-of-its-kind benchmark to discern human-written and LLM-generated science news.

    We demonstrate that Guided Few-shot, a prompting template effectively boosts the performance of LLMs to discern human-written and LLM-generated news. Specifically, GFS boosted the performance of Llama 3, an open-weight LLM, to the performance comparable to commercial LLMs.

2 Related Work

2.1 LLM-based Text Evaluator

The metrics of evaluating the AI-generated text have been studied extensively, e.g., [13, 39]. Traditionally, they rely on token-based metrics such as ROUGE [24], BLEU [33], or BERTScore [44], which are largely based on $n$-gram overlap with reference texts and do not necessarily account for comprehensiveness, accuracy, and readability [35]. LLMs have exhibited a promising capability in reading and comprehension, question answering, text generation, reasoning, and scientific knowledge [7, 28]. A recent study reports >90% accuracy when flagging LinkedIn profiles that are entirely LLM-generated [4]. LLMs have also been explored as evaluators of generated text in machine translation [22], close-ended responses [25], and summarization [17]. Due to the limitations of token-based evaluation metrics, pairwise comparison was proposed as a more reliable way for text evaluation because it was better aligned with human preferences [27], although some studies provide counterexamples [8]. Others have introduced a reader-aware writing assistant that rewrites technical prose via LLMs and evaluated it through human-vs-machine pairwise comparisons [23]. The capability of LLMs to discern human-written and LLM-generated news articles under the pairwise settings remains understudied.

2.2 Science News Datasets

Existing datasets for news generation and summarization, such as CNN/Daily Mail [30], primarily focus on general news [40]. This has led to extensive research in these areas e.g., [37, 40]. However, datasets specifically tailored for scientific domains remain limited. The news summarization dataset, XSUM, introduced by Narayan et al. [31] comprised news articles published on BBC.com, spanning a variety of domains including Technology, Science, and Health.

Several datasets have been developed for scientific text simplification and summarization [6, 16, 21]. A recent study presented the SciNews dataset, which include a collection of science news articles and their corresponding research papers [34]. The news articles are compiled from ScienceX, a network of websites dedicated to science, technology, and medical news. However, this dataset is not manually annotated to topically align text with scientific papers. Our dataset, although smaller, contains manually annotated text to ensure the text is topically aligned with scientific papers.

3 Data

We construct a new dataset called SANews containing 362 news articles collected from ScienceAlert, [36] a news outlet that independently publishes articles featuring scientific research. Each news article is manually annotated into one or multiple snippets each describing a scientific paper. Each snippet is also accompanied by an LLM-generated news article, using a basic prompt grounded on the paper abstract. All are encapsulated in a triplet containing

(human-written news, paper abstract, LLM-generated news).
| Category | #Papers | #News Articles | |---------------|---------|----------------| | Nature | 49 | 44 | | Uncategorized | 70 | 67 | | Tech | 42 | 41 | | Physics | 19 | 19 | | Health | 59 | 56 | | Space | 34 | 34 | | Humans | 51 | 57 | | Environment | 37 | 35 | | Society | 1 | 1 | | **Total** | **362** | **344** |

Table 1: Subject categories of news articles in SANews.

Figure 2: Word count distribution for annotated snippets. The red line shows the mean. The summary statistics are presented in the upper right corner. The histogram was smoothed using Kernel Density Estimate [9].

3.1 Data Collection and Pre-Processing

The raw data was collected using a focused web crawler, which automatically downloaded 23,674 HTML webpages that contain science news articles from ScienceAlert.com dated from 2014 to 2022. The seed URLs were obtained by looping through the links on the website's archive pages. We then automatically parsed the HTML pages and extracted the news text and metadata, including title, category, publication dates, and reference links.

Because not all ScienceAlert news articles contain URLs referencing scientific papers, we identified 8,743 news articles that contain a total of 16,278 unique URLs linking to scientific papers. A URL was identified as linking to a scientific paper if it was from a publisher's domain and/or contained a DOI. We compiled a list consisting of 59 publisher domains mostly adopted from the CiteSeerX crawler's blacklist [42].

3.2 Annotation

We selected a random subset of 748 news articles and manually obtained the 865 paper abstracts that were referenced by these news articles from the publisher's websites. Next, we annotated news text snippets that directly describes each scientific paper.

The annotations were performed independently by two graduate students. A snippet may contain a paper's background, motivation, methods, data, claims, conclusions, quotes, and remarks by authors and journalists. After the annotation, the final dataset was selected by calculating the similarity scores of two annotators using the longest common subsequence algorithm. Annotations with a similarity score greater than or equal to 0.85 were incorporated into the final dataset. Through this effort, we compiled SANews that contains news text that was precisely aligned with scientific papers. The distributions of papers across eight categories are shown in Table 1. On average, each text snippet contains 532 words.

Given the paper abstract, we adopted GPT-3.5 to draft a news article as a "journalist" based on each abstract. Although new GPT versions may be more powerful in text generation, the goals are to test if LLMs can beat humans on the discerning tasks and demonstrate the proposed method outperforms the baseline methods, as opposed to establishing a new state-of-the-art for text detection. Therefore, we consistently use news generated by GPT-3.5. We will demonstrate that even news generated by GPT-3.5 presented a challenge for many LLMs. Future research will delve into more sophisticated methods for generating science news. We randomly split the dataset into 60 training and 302 test samples.

4 Methods

4.1 Guided Few-shot

Inspired by Chain-of-Thought (CoT) [41], we propose a simple yet effective approach, named Guided Few-shot (GFS), which distinguishes human-written news from LLM-generated news under both pairwise and single-item settings. The main difference between CoT and GFS is that GFS's reasoning is based on a set of unordered guides instead of stepwise instructions that must follow a certain order. The GFS method is comprised of the following components:


    System prompt: Assign the LLM a role in this task. In our case, the role is an evaluator who reads and understands scientific papers and related news articles.

    Task instructions: Define the task. In our case, the task is to determine which news is more likely written by humans.

    Examples: Provide examples as triplets.

    Guidance: Provide the guidance of this task in terms characteristics of properties of different types of text. In our case, we provide the characteristics of human-written articles and AI-generated content.

    Task assignment: Provide the task triplet and request the LLM to accomplish the task.

Developing the proper guidance is essential for the success of the task. Instead of simply handcraft guidance based on our intuition, we attempt to generate complete and accurate guidance by fusing the guidance provided by the LLM and the guidance from the humans. We first prompt an LLM to accomplish a small set of seed tasks and provide explanations of the rubrics of decisions. The news of the seed tasks are carefully chosen to represent news articles in diverse domains. Then, the rubrics are inspected and revised by humans to generate the guidance. The prompting template to accompolish our task is shown in Figure 3.

Figure 3: The GFS prompt template (pairwise setting).

4.2 Baseline Methods

The baseline methods include the following.


    GPTZero. A model designed to discriminate between human-generated and AI-generated text based on perplexity and burstiness.

    Direct. This approach directly compares the human-written and LLM-generated news articles in either pairwise or single-item settings (see below).

    Guidance Only. Requesting the LLM to accomplish the task by providing the LLM with guidance.

    One-shot only. Request the LLM to accomplish the task with only one example provided (no guidance) one-shot learning without providing guidance.

5 Experiments

We explore whether LLMs' performance in discerning between human-written and LLM-generated science news against the SANews dataset. We investigate both pairwise and single-item settings. In the pairwise setting, the model is given a pair of (human-written, LLM-generated) news and the task is to predict which article is likely to be written by a human. In the single-item setting, the model is given a single article (human-written or LLM-generated) at a time and decides whether it is written by a human or an LLM.

5.1 LLMs

We explored nine commercial and open-weight LLMs and tested their effectiveness based on SANews. The commercial models include OpenAI's GPT-3.5, GPT-4, and GPT-4o, and Anthropic's Claude-3.5 Sonnet and Claude-3 Opus. Open-weight LLMs have gained popularity in the research community for their free cost and their increasing capabilities for many downstream tasks. We used LLaMA-3 (8b and 70b) instruct variants [14], and Mistral 7b [18].

We deployed pre-trained LLaMA-3 models (8b and 70b) and Mistral 7b on a high-performance computing cluster, utilizing two NVIDIA A100 GPUs, each equipped with 80 GB of memory. After the models were loaded, inference with the LLaMA-3 70b model required 1 GPU hour per experiment, while inference with the LLaMA-3 8b model took 20 minutes per experiment. The inference time for Mistral 7b was comparable to that of LLaMA-3 8b.

5.2 Human Study

We conducted a small-scale experiment to test the discernability of junior graduate students in a U.S. university. Employing graduate students instead of the general public allows us to establish a strong baseline for LLMs. We randomly selected 3 junior graduate students in computer science. Each participant was asked to read 11 pairs of (human-written, LLM-generated) news articles. The order in which each pair was given and the order of news within each pair were randomized. The participant was then asked to judge whether the article was written by humans.

5.3 Evaluation Metric

The evaluation metrics used throughout this study include precision, recall, and F1-score. In the pairwise setting, it is calculated as the number of correct prompt responses divided by the total number of prompt responses. Because GPTZero accepts a single piece of text, its accuracy is calculated the same way as the sequential classification.

| Setting | Model | Direct P | Direct R | Direct F1 | Guidance Only P | Guidance Only R | Guidance Only F1 | One-shot Only P | One-shot Only R | One-shot Only F1 | GFS ($n=1$) P | GFS ($n=1$) R | GFS ($n=1$) F1 | |---|---|---|---|---|---|---|---|---|---|---|---|---|---| | Pairwise | GPT-3.5 | 0.085 | 0.080 | 0.082 | 0.223 | 0.315 | 0.249 | 0.166 | 0.279 | 0.207 | 0.278 | 0.285 | 0.278 | | Pairwise | GPT-4 | 0.186 | 0.233 | 0.207 | 0.480 | 0.523 | 0.378 | 0.699 | 0.699 | 0.699 | 0.932 | 0.931 | 0.930 | | Pairwise | GPT-4o | 0.213 | 0.285 | 0.235 | 0.935 | 0.897 | **0.912** | 0.771 | 0.769 | **0.769** | 0.924 | 0.923 | 0.923 | | Pairwise | C-3 Opus | 0.121 | 0.146 | 0.129 | 0.846 | 0.820 | **0.815** | 0.780 | 0.620 | 0.546 | 1.000 | 1.000 | **1.000** | | Pairwise | C-3.5 Sonnet | 0.207 | 0.212 | 0.206 | 0.793 | 0.676 | 0.637 | 0.993 | 0.993 | **0.993** | 1.000 | 1.000 | **1.000** | | Pairwise | L-3 8b | 0.226 | 0.974 | 0.367 | 0.778 | 0.778 | 0.777 | 0.661 | 0.472 | 0.346 | 0.732 | 0.607 | 0.565 | | Pairwise | L-3 70b | 0.219 | 0.404 | 0.284 | 0.857 | 0.798 | 0.791 | 0.773 | 0.721 | 0.710 | 0.984 | 0.983 | **0.983** | | Pairwise | Mistral 7b | 0.979 | 0.510 | **0.670** | 1.000 | 0.482 | 0.651 | 0.951 | 0.502 | 0.657 | 1.000 | 0.4917 | 0.6592 | | Single-item | GPT-3.5 | 0.750 | 0.500 | 0.336 | 0.250 | 0.500 | 0.333 | 0.405 | 0.469 | 0.362 | 0.811 | 0.727 | 0.707 | | Single-item | GPT-4 | 0.507 | 0.505 | 0.465 | 0.248 | 0.488 | 0.330 | 0.640 | 0.616 | 0.600 | 0.697 | 0.510 | 0.359 | | Single-item | GPT-4o | 0.527 | 0.518 | 0.477 | 0.250 | 0.488 | 0.330 | 0.674 | 0.669 | 0.666 | 0.708 | 0.515 | 0.368 | | Single-item | C-3 Opus | 0.622 | 0.540 | 0.447 | 0.795 | 0.745 | **0.734** | 0.879 | 0.840 | 0.836 | 0.946 | 0.940 | 0.934 | | Single-item | C-3.5 Sonnet | 0.825 | 0.730 | **0.709** | 0.980 | 0.980 | **0.980** | 0.995 | 0.995 | **0.995** | 0.998 | 0.998 | **0.998** | | Single-item | L-3 8b | 0.147 | 0.204 | 0.170 | 0.615 | 0.523 | 0.404 | 0.554 | 0.431 | 0.484 | 0.232 | 0.326 | 0.271 | | Single-item | L-3 70b | 0.250 | 0.500 | 0.333 | 0.250 | 0.500 | 0.333 | 0.848 | 0.782 | 0.770 | 0.990 | 0.990 | **0.990** | | Single-item | Mistral 7b | 0.751 | 0.505 | 0.344 | 0.751 | 0.505 | 0.344 | 0.651 | 0.588 | 0.540 | 0.754 | 0.517 | 0.369 |

Table 2: Top: Comparisons of the weighted average precision ($P$), recall ($R$), and $F1$-scores between GFS ($n = 1$) and baselines in the pairwise setting. Most evaluations were based on 302 testing samples in SANews. Due to the limited budget, we used 100 random samples for Claude-3 (C-3) Opus. Human's performance is ($P$,$R$,$F1$)=(1.0, 0.97, 0.98) (Section 5.2). The top two performances for each method are highlighted in bold. Bottom: Comparisons of weighted average precision ($P$), recall ($R$), and $F1$-scores between GFS ($n - 1$) and baselines in the single-item setting. Most evaluations were based on 604 testing samples. Due to the limited budget, we have only conducted the single-item classification using 200 samples for Claude-3 Opus. For comparison, the weighted ($P$, $R$, $F1$) for GPTZero are (0.988, 0.969, 0.978). C-3=Claude-3. L-3=Llama-3.

5.4 Results

Human readers, without guidance or examples, achieved an overall F1 of 0.98, which is on par with GPTZero (0.978). In contrast, Table 2 shows that, under the zero-shot conditions (Direct), LLMs only achieved and F1 of 0.082–0.670 under the pairwise setting and 0.170–0.709 under the single-item setting, which indicates that although the discerning task is relatively straightforward for human readers in our study, LLMs faced considerable challenges to the task.

Using GFS ($n = 1$), three LLMs surpassed human performance. Specifically, LLaMA-3 70b achieved an F1 of 0.983 under the pairwise setting and 0.990 under the single-item setting. Both Claude 3 and Claude 3.5 attained a perfect F1 under the pairwise setting.

Table 2 shows that when given a single example (One-shot only), most LLMs struggled except for Claude-3.5 Sonnet, which achieves and F1 of 0.993 under the pairwise setting and 0.995 under the single-item setting. Similarly, when given the guidance only, most LLMs struggled except for GPT-4o, which achieves an F1 of 0.912 under the pairwise setting and Claude-3.5 Sonnet, which achieves an F1 of 0.980 under the single-item setting.

The results indicate it was possible to use the GFS method to boost the performance of LLMs to discern human-written and LLM-generated science news. Open-weight LLMs such as LLaMA-3 70b may be suitable for this task as the performance has reached or exceeded the human level and is on par with the best commercial LLMs, as shown in Table 2.

6 Discussion and Conclusions

One limitation of our work is that the LLM-based method to generate science news is relatively simple. Human-written science news is usually based on a broad scope of materials such as the full text and interviews, but LLM-generated news is generated based on the abstract. More sophisticated news generation models can be explored to make the discerning task more challenging.

In conclusion, we investigated the question of whether LLMs can beat human readers in discerning between human-written and LLM-generated science news. Tested against SANews, a novel dataset we found that under both pairwise and single-item settings, by providing proper guidance and a single example, it was possible to boost the performance of commercial and open-weight LLMs to achieve or exceed the performance of human readers and non-LLM methods.

References


    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023).

    Kholoud Aldous, Joni Salminen, Ali Farooq, Soon-gyo Jung, and Bernard Jansen. 2024. Using ChatGPT in content marketing: enhancing users' social media engagement in cross-platform content creation through generative AI. In Proceedings of the 35th ACM Conference on Hypertext and Social Media. 376–383.

    Anthropic. 2023. Claude. https://www.anthropic.com/news/claude-3-family. Accessed: 2024-07-11.

    Navid Ayoobi, Sadat Shahriar, and Arjun Mukherjee. 2023. The looming threat of fake and llm-generated linkedin profiles: Challenges and opportunities for detection and prevention. In Proceedings of the 34th ACM Conference on Hypertext and Social Media. 1–10.

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. arXiv:2005.14165 [cs.CL] https://arxiv.org/abs/2005.14165

    Muthu Kumar Chandrasekaran, Anita de Waard, Guy Feigenblat, Dayne Freitag, Tirthankar Ghosal, Eduard Hovy, Petr Knoth, David Konopnicki, Philipp Mayr, Robert M. Patton, and Michal Shmueli-Scheuer (Eds.). 2020. Proceedings of the First Workshop on Scholarly Document Processing. Association for Computational Linguistics, Online. https://aclanthology.org/2020.sdp-1.0

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. 2024. A Survey on Evaluation of Large Language Models. ACM Trans. Intell. Syst. Technol. 15, 3, Article 39 (March 2024), 45 pages. doi:10.1145/3641289

    Yi Chen, Rui Wang, Haiyun Jiang, Shuming Shi, and Ruifeng Xu. 2023. Exploring the use of large language models for reference-free text quality evaluation: An empirical study. arXiv preprint arXiv:2304.00723 (2023).

    Yen-Chi Chen. 2017. A tutorial on kernel density estimation and recent advances. Biostatistics & Epidemiology 1, 1 (2017), 161–187.

    Matthew Conlen. 2024. We Built a News Site Powered by LLMs and Public Data: Here's What We Learned. https://generative-ai-newsroom.com/we-built-a-news-site-powered-by-llms-and-public-data-heres-what-we-learned-aba6c52a7ee4.

    Xiao Fang, Shangkun Che, Minjia Mao, Hongzhe Zhang, Ming Zhao, and Xiaohang Zhao. 2024. Bias of AI-generated content: an examination of news produced by large language models. Scientific Reports 14, 1 (2024), 5224. doi:10.1038/s41598-024-55686-2

    Cary Funk, Jeffrey Gottfried, and Amy Mitchell. 2017. Science News and Information Today. https://www.pewresearch.org/science/2017/09/20/science-news-and-information-today/.

    Mingqi Gao, Xinyu Hu, Jie Ruan, Xiao Pu, and Xiaojun Wan. 2024. Llm-based nlg evaluation: Current status and challenges. arXiv preprint arXiv:2402.01383 (2024).

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aurelien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Roziere, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, Danny Wyatt, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Francisco Guzmán, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Govind Thattai, Graeme Nail, Gregoire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel Kloumann, Ishan Misra, Ivan Evtimov, Jack Zhang, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Karthik Prasad, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, Khalid El-Arini, Krithika Iyer, Kshitiz Malik, Kuenley Chiu, Kunal Bhalla, Kushal Lakhotia, Lauren Rantala-Yeary, Laurens van der Maaten, Lawrence Chen, Liang Tan, Liz Jenkins, Louis Martin, Lovish Madaan, Lubo Malo, Lukas Blecher, Lukas Landzaat, Luke de Oliveira, Madeline Muzzi, Mahesh Pasupuleti, Mannat Singh, Manohar Paluri, Marcin Kardas, Maria Tsimpoukelli, Mathew Oldham, Mathieu Rita, Maya Pavlova, Melanie Kambadur, Mike Lewis, Min Si, Mitesh Kumar Singh, Mona Hassan, Naman Goyal, Narjes Torabi, Nikolay Bashlykov, Nikolay Bogoychev, Niladri Chatterji, Ning Zhang, Olivier Duchenne, Onur Çelebi, Patrick Alrassy, Pengchuan Zhang, Pengwei Li, Petar Vasic, Peter Weng, Prajjwal Bhargava, Pratik Dubal, Praveen Krishnan, Punit Singh Koura, Puxin Xu, Qing He, Qingxiao Dong, Ragavan Srinivasan, Raj Ganapathy, Ramon Calderer, Ricardo Silveira Cabral, Robert Stojnic, Roberta Raileanu, Rohan Maheswari, Rohit Girdhar, Rohit Patel, Romain Sauvestre, Ronnie Polidoro, Roshan Sumbaly, Ross Taylor, Ruan Silva, Rui Hou, Rui Wang, Saghar Hosseini, Sahana Chennabasappa, Sanjay Singh, Sean Bell, Seohyun Sonia Kim, Sergey Edunov, Shaoliang Nie, Sharan Narang, Sharath Raparthy, Sheng Shen, Shengye Wan, Shruti Bhosale, Shun Zhang, Simon Vandenhende, Soumya Batra, Spencer Whitman, Sten Sootla, Stephane Collot, Suchin Gururangan, Sydney Borodinsky, Tamar Herman, Tara Fowler, Tarek Sheasha, Thomas Georgiou, Thomas Scialom, Tobias Speckbacher, Todor Mihaylov, Tong Xiao, Ujjwal Karn, Vedanuj Goswami, Vibhor Gupta, Vignesh Ramanathan, Viktor Kerkez, Vincent Gonguet, Virginie Do, Vish Vogeti, Vítor Albiero, Vladan Petrovic, Weiwei Chu, Wenhan Xiong, Wenyin Fu, Whitney Meers, Xavier Martinet, Xiaodong Wang, Xiaofang Wang, Xiaoqing Ellen Tan, Xide Xia, Xinfeng Xie, Xuchao Jia, Xuewei Wang, Yaelle Goldschlag, Yashesh Gaur, Yasmine Babaei, Yi Wen, Yiwen Song, Yuchen Zhang, Yue Li, Yuning Mao, Zacharie Delpierre Coudert, Zheng Yan, Zhengxing Chen, Zoe Papakipos, Aaditya Singh, Aayushi Srivastava, Abha Jain, Adam Kelsey, Adam Shajnfeld, Adithya Gangidi, Adolfo Victoria, Ahuva Goldstand, Ajay Menon, Ajay Sharma, Alex Boesenberg, Alexei Baevski, Allie Feinstein, Amanda Kallet, Amit Sangani, Amos Teo, Anam Yunus, Andrei Lupu, Andres Alvarado, Andrew Caples, Andrew Gu, Andrew Ho, Andrew Poulton, Andrew Ryan, Ankit Ramchandani, Annie Dong, Annie Franco, Anuj Goyal, Aparajita Saraf, Arkabandhu Chowdhury, Ashley Gabriel, Ashwin Bharambe, Assaf Eisenman, Azadeh Yazdan, Beau James, Ben Maurer, Benjamin Leonhardi, Bernie Huang, Beth Loyd, Beto De Paola, Bhargavi Paranjape, Bing Liu, Bo Wu, Boyu Ni, Braden Hancock, Bram Wasti, Brandon Spence, Brani Stojkovic, Brian Gamido, Britt Montalvo, Carl Parker, Carly Burton, Catalina Mejia, Ce Liu, Changhan Wang, Changkyu Kim, Chao Zhou, Chester Hu, Ching-Hsiang Chu, Chris Cai, Chris Tindal, Christoph Feichtenhofer, Cynthia Gao, Damon Civin, Dana Beaty, Daniel Kreymer, Daniel Li, David Adkins, David Xu, Davide Testuggine, Delia David, Devi Parikh, Diana Liskovich, Didem Foss, Dingkang Wang, Duc Le, Dustin Holland, Edward Dowling, Eissa Jamil, Elaine Montgomery, Eleonora Presani, Emily Hahn, Emily Wood, Eric-Tuan Le, Erik Brinkman, Esteban Arcaute, Evan Dunbar, Evan Smothers, Fei Sun, Felix Kreuk, Feng Tian, Filippos Kokkinos, Firat Ozgenel, Francesco Caggioni, Frank Kanayet, Frank Seide, Gabriela Medina Florez, Gabriella Schwarz, Gada Badeer, Georgia Swee, Gil Halpern, Grant Herman, Grigory Sizov, Guangyi, Zhang, Guna Lakshminarayanan, Hakan Inan, Hamid Shojanazeri, Han Zou, Hannah Wang, Hanwen Zha, Haroun Habeeb, Harrison Rudolph, Helen Suk, Henry Aspegren, Hunter Goldman, Hongyuan Zhan, Ibrahim Damlaj, Igor Molybog, Igor Tufanov, Ilias Leontiadis, Irina-Elena Veliche, Itai Gat, Jake Weissman, James Geboski, James Kohli, Janice Lam, Japhet Asher, Jean-Baptiste Gaya, Jeff Marcus, Jeff Tang, Jennifer Chan, Jenny Zhen, Jeremy Reizenstein, Jeremy Teboul, Jessica Zhong, Jian Jin, Jingyi Yang, Joe Cummings, Jon Carvill, Jon Shepard, Jonathan McPhie, Jonathan Torres, Josh Ginsburg, Junjie Wang, Kai Wu, Kam Hou U, Karan Saxena, Kartikay Khandelwal, Katayoun Zand, Kathy Matosich, Kaushik Veeraraghavan, Kelly Michelena, Keqian Li, Kiran Jagadeesh, Kun Huang, Kunal Chawla, Kyle Huang, Lailin Chen, Lakshya Garg, Lavender A, Leandro Silva, Lee Bell, Lei Zhang, Liangpeng Guo, Licheng Yu, Liron Moshkovich, Luca Wehrstedt, Madian Khabsa, Manav Avalani, Manish Bhatt, Martynas Mankus, Matan Hasson, Matthew Lennie, Matthias Reso, Maxim Groshev, Maxim Naumov, Maya Lathi, Meghan Keneally, Miao Liu, Michael L. Seltzer, Michal Valko, Michelle Restrepo, Mihir Patel, Mik Vyatskov, Mikayel Samvelyan, Mike Clark, Mike Macey, Mike Wang, Miquel Jubert Hermoso, Mo Metanat, Mohammad Rastegari, Munish Bansal, Nandhini Santhanam, Natascha Parks, Natasha White, Navyata Bawa, Nayan Singhal, Nick Egebo, Nicolas Usunier, Nikhil Mehta, Nikolay Pavlovich Laptev, Ning Dong, Norman Cheng, Oleg Chernoguz, Olivia Hart, Omkar Salpekar, Ozlem Kalinli, Parkin Kent, Parth Parekh, Paul Saab, Pavan Balaji, Pedro Rittner, Philip Bontrager, Pierre Roux, Piotr Dollar, Polina Zvyagina, Prashant Ratanchandani, Pritish Yuvraj, Qian Liang, Rachad Alao, Rachel Rodriguez, Rafi Ayub, Raghotham Murthy, Raghu Nayani, Rahul Mitra, Rangaprabhu Parthasarathy, Raymond Li, Rebekkah Hogan, Robin Battey, Rocky Wang, Russ Howes, Ruty Rinott, Sachin Mehta, Sachin Siby, Sai Jayesh Bondu, Samyak Datta, Sara Chugh, Sara Hunt, Sargun Dhillon, Sasha Sidorov, Satadru Pan, Saurabh Mahajan, Saurabh Verma, Seiji Yamamoto, Sharadh Ramaswamy, Shaun Lindsay, Shaun Lindsay, Sheng Feng, Shenghao Lin, Shengxin Cindy Zha, Shishir Patil, Shiva Shankar, Shuqiang Zhang, Shuqiang Zhang, Sinong Wang, Sneha Agarwal, Soji Sajuyigbe, Soumith Chintala, Stephanie Max, Stephen Chen, Steve Kehoe, Steve Satterfield, Sudarshan Govindaprasad, Sumit Gupta, Summer Deng, Sungmin Cho, Sunny Virk, Suraj Subramanian, Sy Choudhury, Sydney Goldman, Tal Remez, Tamar Glaser, Tamara Best, Thilo Koehler, Thomas Robinson, Tianhe Li, Tianjun Zhang, Tim Matthews, Timothy Chou, Tzook Shaked, Varun Vontimitta, Victoria Ajayi, Victoria Montanez, Vijai Mohan, Vinay Satish Kumar, Vishal Mangla, Vlad Ionescu, Vlad Poenaru, Vlad Tiberiu Mihailescu, Vladimir Ivanov, Wei Li, Wenchen Wang, Wenwen Jiang, Wes Bouaziz, Will Constable, Xiaocheng Tang, Xiaojian Wu, Xiaolan Wang, Xilun Wu, Xinbo Gao, Yaniv Kleinman, Yanjun Chen, Ye Hu, Ye Jia, Ye Qi, Yenda Li, Yilin Zhang, Ying Zhang, Yossi Adi, Youngjin Nam, Yu, Wang, Yu Zhao, Yuchen Hao, Yundi Qian, Yunlu Li, Yuzi He, Zach Rait, Zachary DeVito, Zef Rosnbrick, Zhaoduo Wen, Zhenyu Yang, Zhiwei Zhao, and Zhiyu Ma. 2024. The Llama 3 Herd of Models. arXiv:2407.21783 [cs.AI] https://arxiv.org/abs/2407.21783

    Valentin Grimm and Jessica Rubart. 2024. Authoring Educational Hypercomics assisted by Large Language Models. In Proceedings of the 35th ACM Conference on Hypertext and Social Media. 88–97.

    Yue Guo, Wei Qiu, Yizhong Wang, and Trevor Cohen. 2021. Automated lay language summarization of biomedical scientific reviews. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 160–168.

    Sameer Jain, Vaishakh Keshava, Swarnashree Mysore Sathyendra, Patrick Fernandes, Pengfei Liu, Graham Neubig, and Chunting Zhou. 2023. Multi-dimensional evaluation of text summarization with in-context learning. arXiv preprint arXiv:2306.01200 (2023).

    Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023).

    Kai Jiang, Qilai Zhang, Dongsheng Guo, Dengrong Huang, Sijia Zhang, Zizhong Wei, Fanggang Ning, and Rui Li. 2024. AI-Generated News Articles Based on Large Language Models. In Proceedings of the 2023 International Conference on Artificial Intelligence, Systems and Network Security (Mianyang, China) (AISNS '23). Association for Computing Machinery, New York, NY, USA, 82–87. doi:10.1145/3661638.3661654

    Katikapalli Subramanyam Kalyan. 2023. A survey of GPT-3 family large language models including ChatGPT and GPT-4. Natural Language Processing Journal (2023), 100048.

    Yea-Seul Kim, Jessica Hullman, Matthew Burgess, and Eytan Adar. 2016. Simplescience: Lexical simplification of scientific terminology. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. 1066–1071.

    Tom Kocmi and Christian Federmann. 2023. Large Language Models Are State-of-the-Art Evaluators of Translation Quality. In Proceedings of the 24th Annual Conference of the European Association for Machine Translation, Mary Nurminen, Judith Brenner, Maarit Koponen, Sirkku Latomaa, Mikhail Mikhailov, Frederike Schierl, Tharindu Ranasinghe, Eva Vanmassenhove, Sergi Alvarez Vidal, Nora Aranberri, Mara Nunziatini, Carla Parra Escartín, Mikel Forcada, Maja Popovic, Carolina Scarton, and Helena Moniz (Eds.). European Association for Machine Translation, Tampere, Finland, 193–203. https://aclanthology.org/2023.eamt-1.19

    Ge Li, Danai Vachtsevanou, Jérémy Lemée, Simon Mayer, and Jannis Strecker. 2024. Reader-aware Writing Assistance through Reader Profiles. In Proceedings of the 35th ACM Conference on Hypertext and Social Media. 344–350.

    Chin-Yew Lin. 2004. ROUGE: A Package for Automatic Evaluation of Summaries. In Text Summarization Branches Out. Association for Computational Linguistics, Barcelona, Spain, 74–81. https://aclanthology.org/W04-1013

    Yongkang Liu, Shi Feng, Daling Wang, Yifei Zhang, and Hinrich Schütze. 2023. Evaluate What You Can't Evaluate: Unassessable Quality for Generated Response. arXiv preprint arXiv:2305.14658 (2023).

    Yiheng Liu, Tianle Han, Siyuan Ma, Jiayue Zhang, Yuanyuan Yang, Jiaming Tian, Hao He, Antong Li, Mengshen He, Zhengliang Liu, et al. 2023. Summary of chatGPT/GPT-4 research and perspective towards the future of large language models. arXiv. arXiv preprint arXiv:2304.01852 (2023).

    Adian Liusie, Potsawee Manakul, and Mark Gales. 2024. LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), Yvette Graham and Matthew Purver (Eds.). Association for Computational Linguistics, St. Julian's, Malta, 139–151. https://aclanthology.org/2024.eacl-long.8

    Rui Mao, Guanyi Chen, Xulang Zhang, Frank Guerin, and Erik Cambria. 2023. GPTEval: A survey on assessments of ChatGPT and GPT-4. URL https://arxiv.org/abs/2308.12488 (2023).

    Bonan Min, Hayley Ross, Elior Sulem, Amir Pouran Ben Veyseh, Thien Huu Nguyen, Oscar Sainz, Eneko Agirre, Ilana Heinz, and Dan Roth. 2021. Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey. arXiv:2111.01243 [cs.CL]

    Ramesh Nallapati, Bowen Zhou, Caglar Gulcehre, Bing Xiang, et al. 2016. Abstractive text summarization using sequence-to-sequence rnns and beyond. arXiv preprint arXiv:1602.06023 (2016).

    Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun'ichi Tsujii (Eds.). Association for Computational Linguistics, Brussels, Belgium, 1797–1807. doi:10.18653/v1/D18-1206

    Originality.AI. 2024. Originality: AI Content Detection and More For Serious Web Publishers. https://originality.ai/.

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, July 6-12, 2002, Philadelphia, PA, USA. 311–318. http://www.aclweb.org/anthology/P02-1040.pdf

    Dongqi Pu, Yifan Wang, Jia Loy, and Vera Demberg. 2024. SciNews: From Scholarly Complexities to Public Narratives–A Dataset for Scientific News Report Generation. arXiv preprint arXiv:2403.17768 (2024).

    Ehud Reiter and Anja Belz. 2009. An investigation into the validity of some metrics for automatically evaluating natural language generation systems. Computational Linguistics 35, 4 (2009), 529–558.

    ScienceAlert Pty Ltd. 2024. ScienceAlert. https://www.sciencealert.com/.

    Grishma Sharma and Deepak Sharma. 2022. Automatic text summarization methods: A comprehensive review. SN Computer Science 4, 1 (2022), 33.

    Shehul Singh and Nehul Singh. 2023. GPT-3.5 vs. GPT-4, Unveiling OpenAI's Latest Breakthrough in Language Models. (11 2023). doi:10.36227/techrxiv.24486214.v1

    Ruixiang Tang, Yu-Neng Chuang, and Xia Hu. 2024. The science of detecting llm-generated text. Commun. ACM 67, 4 (2024), 50–59.

    Xuezhi Wang and Cong Yu. 2019. Summarizing news articles using question-and-answer pairs via learning. In The Semantic Web–ISWC 2019: 18th International Semantic Web Conference, Auckland, New Zealand, October 26–30, 2019, Proceedings, Part I 18. Springer, 698–715.

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh (Eds.). http://papers.nips.cc/paper_files/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract-Conference.html

    Jian Wu, Pradeep Teregowda, Juan Pablo Fernández Ramírez, Prasenjit Mitra, Shuyi Zheng, and C. Lee Giles. 2012. The evolution of a crawling strategy for an academic document search engine: whitelists and blacklists. In Proceedings of the 4th Annual ACM Web Science Conference (Evanston, Illinois) (WebSci '12). Association for Computing Machinery, New York, NY, USA, 340–343. doi:10.1145/2380718.2380762

    ZeroGPT. 2024. Trusted GPT-4, ChatGPT and AI Detector tool by ZeroGPT. https://www.zerogpt.com/.

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating Text Generation with BERT. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net. https://openreview.net/forum?id=SkeHuCVFDr

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime