Study Shows Government Media Shapes AI Answers

New research indicates that the language used to query AI models can significantly alter the political tone of the responses, revealing deep structural biases in how these systems learn from the internet.
Researchers at the University of Oregon have identified a direct link between state-controlled media landscapes and the behavior of large language models. The study suggests that AI systems do not learn from a neutral pool of information but rather absorb the specific rhetorical styles and biases present in the online content they are trained on. This means that the political tone of an AI's response can shift depending on the language in which the user asks their question.
The implications for users are significant. If a model is trained heavily on media from a specific government, it may produce answers that are more favorable to that government when queried in the local language compared to English. This creates a scenario where the same factual question yields different political leanings based solely on the linguistic interface, highlighting a trade-off between global consistency and local data availability.
Language Drives Political Bias
The core finding of the research is that institutional influence leaves detectable traces in AI output. When the team tested 37 countries, they found that models portrayed governments from nations with strong media control more favorably in the local language than in English. For countries like Turkmenistan and Vietnam, the local language response was more positive over 75% of the time. In contrast, countries with less state media influence, such as Sweden, showed no significant difference between languages.
This discrepancy arises because AI models are trained on vast amounts of web data. In countries where the government heavily shapes the media environment, the training data reflects that specific viewpoint. When a user interacts with the model in that local language, the system relies more on those specific data patterns, resulting in a biased output. The catch for developers and users is that this bias is not always visible in English-centric testing, masking the underlying structural inequality in the model's knowledge base.
State Media Dominates Training Data
To understand the source of this bias, the team analyzed open-source data repositories like Common Crawl, which are widely used for training commercial models. They discovered a substantial overlap between state-coordinated media phrasing and the training documents. In Chinese-language datasets, 3.1 million documents contained phrasing that matched state-controlled sources. This volume was more than 40 times higher than the representation of Chinese Wikipedia articles in the same dataset.
The study notes that while most AI companies do not disclose their exact training data, the prevalence of state media in public repositories suggests a significant input channel. For documents specifically mentioning political leaders, the share of text with overlapping state phrasing rose as high as 23%. This indicates that for certain topics, the model's understanding is heavily weighted toward the official narrative of the state, rather than a diverse array of global perspectives.
Testing Models With Aligned Data
To confirm that this data overlap actually changes model behavior, the researchers conducted experiments with smaller, open-source AI models. They added state-coordinated news content to the training data and observed the results. The models became significantly more likely to produce pro-government answers, particularly when the questions were asked in Chinese rather than English. This demonstrates that the influence is not just a correlation but a causal mechanism driven by the training process.
The researchers applied these findings to commercial models as well. When probing major chatbots, they found that questions about China yielded more favorable answers in Chinese than in English. This reinforces the idea that the language of the prompt acts as a filter, activating different subsets of the model's training data. For users, this means that seeking objective political analysis from AI may require careful consideration of the language used, as the underlying data may carry heavy institutional biases that are not immediately apparent.






