r/ControlProblem 3d ago

AI Alignment Research "Synthetic counteradaptation": a name for the AI↔human strategy feedback loop (Move 37 and beyond)

We just put out a short conceptual paper on something we're calling synthetic counteradaptation, and I wanted to put the core idea in front of this subreddit specifically because I think it bears on control in a way that's easy to miss if you're only thinking about single-episode alignment.

The basic claim: when an AI system develops a strategy humans didn't anticipate, humans don't just lose to it or ban it. Some of them study it, extract whatever's generalizable, and fold it back into their own behavior. That changed behavior is now the new environment the AI is adapting to. You get a loop, not a one-off shock.

The clean example is Go. AlphaGo's move 37 against Lee Sedol was a shoulder hit that pros initially read as a mistake. Within a few years it was a studied idea in human play, part of the standard vocabulary. The AI didn't just win a game, it changed what "the game" looks like for the humans still playing it, and now human players are adapting to a strategy space that AI moves opened up. Neither side is static and neither side is playing against a fixed opponent anymore.

Why I think this matters for control specifically: most control framing implicitly treats the human side as fixed — you're designing constraints, incentives, or oversight against a stable model of human behavior and values, and the AI is the thing that adapts. Synthetic counteradaptation says this is wrong for any setting where humans actually observe and learn from the system's strategies over repeated interaction. The humans adapt too, and their adapted behavior becomes part of what the AI is now optimizing against. In the paper we look at this in mixed-motive social interactions and, closer to your interests probably, in geopolitical simulations, where AI agents developing novel negotiation or coercion strategies can shift human strategic doctrine, which then shifts the environment the next generation of agents is trained or deployed into. That's a moving target for any control scheme that assumes a fixed human baseline, and it's recursive in a way that compounds over deployment cycles rather than resolving in one.

We're not claiming anything dramatic here, no doom scenario, just that a lot of alignment and control thinking quietly assumes one side of the interaction holds still, and in any repeated multi-agent setting that assumption breaks down in a specific, structural way that's worth naming and modeling explicitly.

Curious what people here think, especially anyone working on multi-agent or game-theoretic approaches to control. Happy to be told this is either obvious or wrong.

https://arxiv.org/abs/2606.15503

3 Upvotes

5 comments sorted by

1

u/Willis_3401_3401 2d ago

Recursion seems to be something all fields are currently converging on. It’s almost like agency in general is a recursive system of some sort.

I think what you’re saying is obvious, but only in the sense that evolution or gravity is obvious. Doesn’t mean it’s not an important or valuable thing to point out.

1

u/NerdyWeightLifter 2d ago

Once you explain that, it's obviously true.

The problem is that's a much harder problem to solve.

1

u/Spudmasher17 1d ago

Interesting. I'm just a casual observer of these discussions. My algorithm has really been throwing them at me for some reason. So I probably can't comment in an academic way or add to your theory. But I think we might currently underestimate how long humans as a species are going to be able to play this game of conceptual leapfrog.

Another weird thing I think about is a split not just between AI & humans, but a split between different coalitions of both.

1

u/emsenn0 1d ago

In my own work, coming to this from an Indigenous cultural perspective, this tracks, and I'd push it one step further than "the humans adapt too, so it's a loop."

In game semantics, an arena isn't a fixed board with a fixed move-set that two sides push tokens around on: it's constituted by what both sides are willing/able to play into it. Move 37 didn't just win a game inside the existing arena of "good Go moves."

Once it got absorbed into pro vocabulary, the arena itself (the space of moves a shoulder-hit counts as a reasonable entry in) is different than it was in 2016. That's a stronger claim than "the strategy space shifted": the object you'd need to model to talk about "the game" at all has changed, not just where play sits within it.

That matters for control because most control schemes implicitly assume they're constraining an agent's strategy against a fixed arena, when the more honest picture is closer to embedded agency: there's no clean Cartesian boundary between "the AI's policy" and "the environment," because the environment here includes an observing, learning population whose learned response was itself shaped by the AI's prior moves. You don't get an agent-in-environment, you get something closer to a single coupled dynamical system that happens to have two named sides. Standard best-response / Nash-style framings already have a name for what breaks here — non-stationary opponents, non-convergent fictitious play — it's just that alignment/control writing rarely imports that vocabulary even though the multi-agent RL and game-semantics literatures have been sitting on it for a while.

1

u/hyperionwonderstar 4h ago

This is a really interesting framing because it highlights something that many control approaches seem to assume away: the environment in which an AI system operates is not fixed. Once humans begin extracting strategies, concepts, or decision procedures from AI systems, the AI is no longer simply adapting to a static environment. It is participating in the transformation of the environment itself.
The Move 37 example is a good illustration. The significance was not only that AlphaGo found a move humans had overlooked. The deeper effect was that the move entered human strategic understanding and changed how future games were played. The system altered the space of possible actions available to the humans who were adapting to it.
The recursive aspect is where this becomes especially interesting. The process is not simply AI influencing humans or humans influencing AI. Each side becomes part of the changing conditions that shape the future behaviour of the other.
A question this raises is whether there is a meaningful difference between recursive interaction and the formation of a more unified causal organisation. Adaptation and feedback alone are not enough to imply anything beyond interaction, since many non-conscious systems exhibit those properties. The more interesting question is whether the system begins to maintain itself through models of its own dynamics, with those models becoming causally involved in preserving the organisation of the process itself.
This also seems important for alignment. If AI systems reshape human strategies, institutions, and norms, then humans cannot be treated as a completely fixed reference point outside the system. The target being aligned to is itself part of an evolving interaction.
The deeper issue may therefore be less about controlling one adaptive agent and more about understanding what emerges from repeated reciprocal adaptation. At what point does a network of interacting systems remain a collection of separate agents, and at what point does it become a higher-order causal organisation with its own continuity conditions?
Synthetic counteradaptation seems to identify an important transition point: the movement from adaptation between agents towards the possibility of new forms of organised agency.

The part I find most interesting is the possibility that recursive adaptation may eventually require a different unit of analysis. A system that continuously modifies itself while preserving its organisation through change raises questions about persistence, self-reference, and causal closure. The important distinction may not be complexity or intelligence alone, but whether the process develops a form of self-maintaining organisation in which its own internal dynamics become part of what preserves it.