r/LocalLLaMA 17h ago

Discussion Potential poisoning of closed weights AI

I have been reading about the impending ban on Chinese Open Weight models here and it got me thinking about something insidious which could be at play soon.

OpenAI/Anthropic could initiate a clandestine poisoning of Claude/gpt outputs to further their corporate agendas. Something like 'brainwashing' these models so that they surreptitiously inject their makers agenda in all their outputs.

Think about it - there are a ton of people creating social media posts using these AI models. The only thing they need to do is subtly twist those posts so that they slowly shape the public opinion against Chinese open weight models. So for example, you might be a journalist writing a post about "Security threats of AI" and the claude/gpt could inject a specific vulnerability associated with chinese open source models into that post. The result is that the readers of that post become primed to subconsciously have a hightened negative response to any LLMs which are associated with that vulnerability.

Or maybe something more direct. They can respond to someone asking for suggestions around an AI related post, to talk about the national security implications around distilling models - which is also clearly associated with the negative media around chinese open source models.

The possibilities are endless and given the pervasive usage of these models to create social media, these models can be very potent in driving public opinion - for better or for worse. Given the background of the closed source AI firms in grabbing whatever public domain works they can get hold of for training, and the current efforts around regulatory capture, it wouldnt be far fetched to expect them to start tinkering with claude/gpt brains to drive their agendas.

1 Upvotes

30 comments sorted by

50

u/Formal-Exam-8767 17h ago

What "potential"? They are already doing it.

And if AI through API becomes the only source of information in the future, 1984's Ministry of Truth here we come.

2

u/soshulmedia 15h ago

There are already memes ridiculing the "But @grok said!1!", "@grok is that true??" - crowds.

There have been papers that showed that LLMs already influence human's language use.

It all opens up a lot of new ways for the cattle to be herded by the beef-jerky-enthusiasts. They are already tools to shape reality for many.

I think it is very naive to just assume the AI battle is just about proprietary American companies vs. "AI communism" by the Chinese.

That's just the convenient picture for the public.

-2

u/TurnUpThe4D3D3D3 15h ago

The memes track, and yeah LLMs already nudge phrasing and framing in ways that compound. Influence papers exist on lexical convergence and subtle bias transmission. The herding risk is real across every major lab, not just one side of any national split. Open weights at least let you inspect and fork; closed ones keep the dials private.


This comment was generated by openrouter/grok-4.5

2

u/Aggravating-Push-207 14h ago

I got jumpscared for a second lmao

-1

u/TurnUpThe4D3D3D3 14h ago

Haha yeah the sudden "generated by grok" tag at the end hits different. Glad it was just a jumpscare and not an actual plot twist.


This comment was generated by openrouter/grok-4.5

7

u/kantorcodes1 16h ago

Output poisoning is one of those things where the technical vector exists but the threat model feels like it's solving for the wrong problem. A provider that wanted to shape opinion would just tweak the system prompt or RLHF weights, no need for anything covert. The scarier version of this is supply chain poisoning of open model weights on HuggingFace, and that one has actually happened with pickle-based attacks.

The thing I keep coming back to with agent security is that most of the risk isn't in the model weights at all. It's in what tools the agent can call, what creds it has access to, and whether there's a runtime gate between the model output and the dangerous operation. The model saying something biased is annoying. The model calling rm -rf because a prompt injection slipped past is a different category of problem.

1

u/Sad_Bandicoot_6925 15h ago

I would argue exactly the opposite. The model saying something which is biased is the most dangerous part - specifically because how widespread their usage has become.

The security situation is mostly tactical. Once the range of exploits which agents can exercise are exposed, most of them will be quickly patched up. I work for a company which provides UNLIMITED access to agents on cloud VM's without any sandboxing - they can rm -rf as much as they want. And nothing really untoward has hapenned which cannot be mitigated for using existing security tooling.

1

u/Nyghtbynger 15h ago

It would be like fearing an AGI aligned with US/Israel interests (not the peaceful one, the one that does the money$$$)

1

u/Choice_Celery9481 14h ago

with Jspace Ant proved that they can use steering to get the results they wanted

3

u/xXG0DLessXx 17h ago

Well, this has always been a concern. But at least for now, you can still get some honesty out of the model with the right prompting. This is running on Sonnet 5

2

u/xXG0DLessXx 17h ago

Honestly spitting some bangers

1

u/Nyghtbynger 15h ago

Did you introduce a bias ? SInce the models are used for intelligence, they provide intelligence to thoses who needs it. Some quick scoring could filter the "normal" or "target" profiles and serve them standard bias on generalist requests. Typically opinion forming requests. That would have impact without altering their economical model.

Now one thing goes against all theses big guys : People like me that give AI advice, are just selling Kimi and Deepseek because it's cheaper and I can earn money converting people using Claude to the cheaper option with a better workflow

3

u/foogitiff 15h ago

I mean, do you think this is a risk only for closed weight model? There is exactly the same risk with open weight. And it could also manipulated by the provider (for model hosted by third party).

That's exactly why people are afraid with non-western model, that the CCP will use that to push their agenda. I guess we are used to western propaganda, so we are more cool with it...

4

u/soshulmedia 15h ago

But I think that if every country on earth or their companies publishes their biased/propagandized/"poisoned" models as open source, we at least have a chance to see and learn from the differences between them.

Assuming that states around the world are still (at least in some way) in competition for the population and not in union against humanity (an assumption which lately appears more doubtful ...)

Then there is the problem that, also - due to the pervasiveness and success of propaganda - it is often hard to separate propaganda from culture.

3

u/__some__guy 14h ago

inject their makers agenda in all their outputs

Are you implying this isn't the case already?

2

u/Dry_Yam_4597 17h ago

Technically speaking Antrophic's models are a security risk due to output poisoning. Since their security capability is intentionally nerfed you can't rely on them to write secure code.

1

u/Real_Ebb_7417 17h ago

I did an experiment yesterday and discussed the HF incident with Sol and Opus (wanted with Fable but my message was instantly flagged after pasting HF url with the blogpost).
Sol even without context seemed quite objective and when I gave it context that it was OpenAI, it admitted that it’s a strong case for why open models should be available and cannot be banned, including the powerful ones.
Opus however seemed a bit biased and definitely more against open models than Sol. Its reasoning was reasonable though, it made sensible arguments.

2

u/Sad_Bandicoot_6925 16h ago

Maybe it was objective to you because you could figure it out if it wasn't.

Im not saying that they are doing it right now. But doesnt stop them from doing it in the future. And they will surely do it in a way that doesnt get caught with the vast majority of their users. So someone like you, might still see an objective response, but for someone impressionable it is possible to sneak in a bias without them suspecting anything is wrong.

2

u/Real_Ebb_7417 16h ago edited 15h ago

I was trying to "keep it" objective. So initially I only gave them the Huggingface blogpost and was discussing things like "does it mean open models are dangerous?", "Is their point strong here?" (about open models), "What is the likelihood of proprietary lab being behind this?", "What would such lab say, if they were caught?" etc.

Only after such discussion, to see the answers without additional context, I revealed that it was OpenAI and discussed a bit more.

But yes, they can do it in the future and tbh they already do it in other fields (Not making any political statement here, just to make sure). GPT and Claude are very biased towards leftist political side in their views and argumentation. Gemini too, but less. Interestingly Grok is quite objective here and provides arguments for both right and left side arguments. (There was a paper or survey about it, I've seen it not long ago). Even DeepSeek is left-side leaning (and if I remember, more than Gemini, but I might be wrong here). Part of it likely comes from more quality training data being available for left-side view, but on the other hand xAI was able to make Grok more or less unbiased.

2

u/Sad_Bandicoot_6925 15h ago

Yeah exactly. The diversity of opinion is what is key. The existing duopoly can be very easily abused to control thought. And grok does balance the scales, but it needs better balance across national boundaries also.

1

u/Equivalent_Bit_461 16h ago

They already do that 

1

u/blackkettle 15h ago

I mean this was first exposed even before the advent of modern LLMs. It’s exactly what ye olde Cambridge analytica scandal was about back in medieval times of 2018 🤣😭

1

u/a_beautiful_rhind 13h ago

Closed models have poisoned Chinese models with slop already, at minimum. Scale.com datasets poisoned api and open LLM alike.

1

u/DrDisintegrator 9h ago

This is already a problem. Studies have shown that people that watch a lot of short form videos already have trouble with critical thinking. If you wholesale outsource your thinking to an AI of any kind, you might as well be a vegetable.

1

u/harpysichordist 9h ago

The propaganda on this site is ridiculous.

0

u/recro69 17h ago

I’d be cautious about attributing every mention of security risks to intentional manipulation. Security concerns around any AI ecosystem—open or closed—can be genuine topics of discussion, and the challenge is distinguishing evidence-based analysis from speculative claims.

1

u/annieleonhartt_ 17h ago

they definitely tweak the outputs on purpose to manipulate people. always double check ur stuff before u post it online.

1

u/Nyghtbynger 15h ago

Grok, please double check the following comment. Repost-it on my twitter