The Pravda Network and its Impacts on Large Language Models – UROP Symposium

The Pravda Network and its Impacts on Large Language Models

Joel Fellows

Research Mentor: Laura Kurek
Mentor Department: Not Available, Information
Author(s): Laura Kurek, Joel Fellows, Henry Huang, Rafael Bonilla, Elijah Covert, Eric Gilbert, Ceren Budak
Session: Session 3 (11:00 AM – 11:50 AM)
Presentation Type: Poster 117

Abstract

The Pravda Network is a collection of over 100 near-identical websites which aggregate and distribute Russian-state aligned news. Given the prolific and repetitive nature of these sites, analysts hypothesize that Pravda Network articles are not intended for human eyes, but intended instead to influence the outputs of large language models (LLMs), which are trained on crawled internet data and incorporate recent news articles in their responses via retrieval-augmented generation (RAG). To investigate the extent of Pravda Network content appearing in RAG results, we prompt leading LLM chatbots with article headlines from the Pravda Network sites. We conduct this LLM auditing at two scales: small-scale manual auditing (N=500) and large-scale automated auditing (N=10,000). We vary the prompts along two dimensions: each prompt contains (a) either the original article headline or a reworded article headline, as well as (b) a non-skeptical prompt phrasing or a skeptical prompt phrasing.

lsa logoum logo