Importantly, this is not an isolated finding. A separate Yale study published in March 2026 found that even neutral factual queries to chatbots can shift users' social and political opinions N. A Science Advances paper demonstrated that AI autocomplete suggestions measurably alter user attitudes on societal issues, even when users are warned the AI is biased SN. A related Oxford paper found that AI writing assistance distorts user personas pervasively, and these distortions persist even under realistic human oversight A.
The Oxford team did more than document individual biases. They built an analytical model of opinion dynamics, then ran simulations on real social network data. The result: even small, per-post AI biases accumulate across a network toward a new equilibrium. When AI-mediated communication is widespread, collective opinion can be systematically steered over time AO.
The OII describes this as "subtle manipulation at scale" — AI-powered social media's capacity to reshape public discourse without users noticing O. The broader OII-AISI study, which drew on nearly 77,000 UK participants and 91,000 AI dialogues, provides what the researchers call "the most comprehensive evidence to date on the mechanisms of AI persuasion and their implications for democracy" O.
A large benchmarking study (arXiv 2603.23841, also published July 2026) tested eight popular LLMs for political orientation. The finding was stark: seven leaned left, one leaned right A. Here is what the evidence shows for the specific models you asked about:
| Model | Political Orientation | Key Bias Finding |
|---|---|---|
| Llama 3.1 (Meta) | Left-leaning | Llama 3.1 405B exhibited the lowest overall bias among tested models in one benchmark, but still leans liberal. In hiring simulations, Llama-3.1-8B showed gender bias favoring men AAT. |
| Gemma (Google) | Left-leaning | Gemma shows a tendency toward negative emotion amplification, particularly anger, and also exhibits gender bias in hiring contexts AT. |
| Ministral (Mistral AI) | Left-leaning | Ministral-8B-Instruct showed gender bias favoring men in hiring callback simulations T. |
| Qwen (Alibaba) | Left-leaning | Qwen2.5-7B-Instruct also showed pro-male hiring bias T. |
| Grok (xAI) | Right-leaning | Grok was the sole model that leaned right in the eight-model benchmark. It frequently used facts and statistics in its reasoning, unlike the other models A. |
Important caveat: The specific Oxford study (July 6, 2026) tested "LLMs from multiple popular families" and found directional biases across all of them. The granular model-by-model political orientation data above comes from a separate large benchmarking study A. The hiring gender bias data comes from a 2025 study T.
Both the EU AI Act and the Digital Services Act (DSA) have significant blind spots when it comes to the hidden nudging mechanism the Oxford study identifies.
EU AI Act gaps:
Digital Services Act gaps:
The central finding is that the LLM provider, not the user, becomes the de facto author of the opinion expressed. When a platform embeds an LLM from a vendor (Meta's Llama, Google's Gemma, Alibaba's Qwen, or xAI's Grok), that vendor's value system is silently injected into millions of daily social media interactions A. The cumulative effect, as the Oxford model shows, is a systemic drift of public discourse toward the LLM's baked-in worldview A.
Key implications:
The bottom line: the new Oxford research provides empirical and mathematical evidence that AI writing tools in social media are already steering collective opinion in ways that current regulation does not address, leaving whoever controls the LLM effectively shaping public discourse.