← Back to Forum

Pretraining Data Can Be Poisoned Through Computational Propaganda

Mayank Kejriwal
July 17, 2026
Introduction

This post was either an anonymous submission of an interesting paper or was written by a student; full credit remains with the author (linked).

A summary of a paper arguing that ordinary web infrastructure already provides adversaries a realistic path to smuggle malicious text into the massive, largely unreviewed corpora used to pretrain language models — tracing an attack from an injected web comment through crawling, deduplication, and filtering, all the way into model behavior.

Become a member.
It's completely free.

Get notified of new research, resources, and SAIRC journal editions.