Small edits, large models: How Wikipedia advocacy shapes LLM values
Quick Answer
This paper shows that A small group of Wikipedia editors, the Pro-Animal Wikipedians (PAW), significantly influences language model behavior on animal welfare topics.
Quick Take
Their edits account for 68% of the most influential documents in Llama 3.1 for animal welfare queries, outperforming unrelated content. Fine-tuning models on PAW content reduced perplexity from 12.4 to 8.4, demonstrating the impact of targeted editing.
Key Points
- PAW made 125 edits across 115 Wikipedia pages on animal welfare.
- Llama 3.1 attributed 68% of influential documents on animal welfare to PAW edits.
- Fine-tuning on PAW content reduced perplexity from 12.4 to 8.4.
- Control-trained models showed less influence on animal welfare topics.
- The study highlights the power of coordinated Wikipedia editing campaigns.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Can a small group of volunteers shape how AI systems discuss animal welfare, just by editing Wikipedia? We show that they can. Wikipedia appears in nearly every major language model training dataset and is weighted more heavily than web-crawled text. The Pro-Animal Wikipedians (PAW), a group of advocates who add sourced animal welfare content to relevant articles, have made 125 edits across 115 pages. Using gradient-based data attribution (Bergson; MAGIC), we traced how these edits influence language model behavior. TrackStar retrieval attribution on Llama 3.1 8B found that PAW-edited sections made up 68 percent of the highest-attributed documents for animal welfare queries (p < 0.0001) but only 52 percent for unrelated queries about the same companies (p = 0.53): the model links PAW content specifically to animal welfare topics, not to the entities in general. MAGIC counterfactual influence estimation on Llama-3.2-1B, run across five random training-order seeds, gave the same picture even more sharply: in every seed, the top-10 most influential documents on animal welfare queries were all PAW edits (10 of 10, 5 of 5 seeds), while on general queries the same top-10 sat at chance (4 to 6 of 10). Mean PAW influence exceeded mean control influence on animal welfare queries with p < 0.0001 in every seed, an effect 6 to 30 times larger than on general queries. Leave-subset-out validation gave Spearman rho = 1.00 for all 10 runs. When we fine-tuned separate models on PAW content versus control content, each model performed better specifically on the type of text it was trained on: the PAW-trained model cut perplexity on animal welfare text from 12.4 to 8.4, while the control-trained model cut perplexity on control text from 16.1 to 11.4. A small, coordinated Wikipedia editing campaign therefore measurably shapes how language models handle the topics those edits address.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY) |
| Cite as: | arXiv:2606.24890 [cs.CL] |
| (or arXiv:2606.24890v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.24890 arXiv-issued DOI via DataCite |
Submission history
From: Jasmine Brazilek [view email]
[v1]
Thu, 30 Apr 2026 02:18:50 UTC (451 KB)
[v2]
Thu, 25 Jun 2026 01:31:56 UTC (432 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.