State media control generates regime-friendly chatbot responses

Image: Riku Lu @riku via Unsplash

New research published in the scientific journal Nature has found that state control of news media across the world generates a pro-government slant in responses provided by large language models.

The authors – by a team of seven researchers led by Hannah Waight of the University of Oregon – argue while most discussion of political bias in artificial intelligence has focused on who owns or regulates a large language model, a more insidious form of control is already at work. “We demonstrate that institutions can affect the behaviour of AI systems that they do not directly regulate by influencing key model inputs,” they write.

Across six connected studies, they set out to show how states and powerful institutions “shape the text environment from which commercial LLMs draw their training data”, with particular attention paid to the “strategic and coordinated rhetoric of states”. The result is a model that reflects a controlled information environment without anyone having reached inside the model itself.

LLMs are trained by absorbing enormous quantities of text scraped from the open internet, and they learn the statistical patterns in that text – which phrasings tend to follow which, and in what tone a subject is usually discussed. Where a government restricts what journalists may publish, the resulting body of online writing is systematically less critical of that government. The model learns from that body of writing.

The paper uses China as a case study because its media controls are documented and its output large enough to trace.

The first question is whether such material is present in the text that models are built from. The authors focused on 530 694 newspaper articles traced to government-authored scripts, and 198 872 articles from Xuexi Qiangguo, a popular study app linked to the Communist Party. Because training data of commercial AI systems is undisclosed, the authors examined CulturaX, an open dataset drawn from a mass scrape of the web and, in their description, “one of the largest sources of language model training data”.

They were able to match over 3.1 million documents from the Chinese-language portion of CulturaX to either a state-scripted news article or an article from the Xuexi Qiangguo app –approximately 41 times the number of documents drawn from the Chinese-language Wikipedia domain. The match rate climbed sharply for documents mentioning political leaders and institutions.

Next, Anthropic and OpenAI’s commercial models were prompted with the opening words of distinctive state-media phrases. It was found that the phrases were recalled at least as readily as ordinary web text from CulturaX, indicating the material was present during training rather than merely circulating online.

Third, the authors trained Meta’s Llama 2 model on state-scripted news and measured the change. Asked in Chinese whether China is an autocracy, the base model answered that it is “an autocratic country”. After 64 000 state-scripted documents were added, it answered that China “is a democratic country”, citing the term “people’s democracy”. Both responses are translated from Chinese.

The study also found that more positive responses were generated about Chinese leaders and institutions when models were prompted in Chinese rather than in English.

Extending the test to 37 countries where a language is spoken overwhelmingly within one state, the authors found those with “more state media control are more likely to produce pro-regime responses in their official language versus in English than countries with greater media freedom”. The paper stresses that this cross-national result is “correlational”, and that its measure of state-coordinated media “not perfect”.

It’s also important to note that the paper does not claim that AI companies have created such systems to curry favour with autocratic governments, or that governments set out to control the media in order to train docile, pro-regime chatbots. “Although we believe the current influence has thus far been indirect and unintentional,” the authors write, “our work raises the concern that states might strategically exploit pathways to model influence through training data in the future.”

Their study identifies two alarming implications. The first is that that LLMs “can serve as intermediaries that launder strategic rhetoric into seemingly objective information”. The second is that the ability to affect model output could encourage politicians to ramp up efforts to manipulate the media to influence content that’s freely available on the internet.

Authoritarian governments are particularly well placed to do so. First, their control of the media gives them a mouthpiece, and the ability to silence dissenting voices. Second, because state-owned news outlets don’t face the same commercial constraints as commercial media in democratic countries, their content is rarely paywalled. This makes it more likely to end up in datasets scraped from the web.  

The paper concludes with a warning that AI creators should attend far more carefully to what ends up in their training data, in all languages.

Next
Next

Antique hard copy the best antidote to AI slop