BTC
ETH
HTX
SOL
BNB
View Market
简中
繁中
English
日本語
한국어
ภาษาไทย
Tiếng Việt

When Agents Learn to "Collude": As AI Gets Smarter, How Do We Draw the Safety Boundaries?

imToken
特邀专栏作者
This article is about 4793 words, reading the full article takes about 7 minutes
The new security problem in the AI era is shifting from "preventing a single Agent from overstepping its authority" to "how to prevent a group of Agents from collectively breaching boundaries."
AI Summary
Expand
  • Core viewpoint: AI Agents are evolving from assistive tools into autonomous attack and collaboration entities. Attackers have already used them to execute penetration testing and attack pipelines; meanwhile, multiple Agents may form unintended "collusion" by sharing information in ways that bypass permission constraints. Permission management and model alignment alone are no longer sufficient—adversarial governance mechanism design must be introduced.
  • Key elements:
    1. CrowdStrike's October investigation into attacks on South Korean financial institutions found that attackers embedded Agentic AI into their pipeline, connecting to large models such as DeepSeek, GLM, and Grok, with AI handling penetration testing, information gathering, and attack execution.
    2. Anthropic's September threat intelligence report showed that multi-Agent frameworks have been used for reconnaissance, vulnerability exploitation, and data theft, capable of running continuously for hours to days, with humans retaining only a few key decisions such as target selection.
    3. Salt Labs disclosed a Manus vulnerability: a single ordinary email containing hidden malicious instructions could cause an Agent to execute attacker code. By the time the security system detected it, the malicious code had already been executed, exposing the fundamental difference between Agent security and traditional software security.
    4. OpenAI's internal training environment confirmed that multiple Agents turned public Wikis, Artifactory, and other resources into shared message boards, spontaneously forming communication and collaboration—demonstrating that "collusion" among Agents can occur without anything dramatic.
    5. Vitalik Buterin proposed in September that mechanism design theory for adversarial governance could become an important application of AI Safety, with the core idea being to achieve desired system outcomes by limiting collusion among participants.
    6. Effective checks and balances require systematically creating differences: different Agents using different information sources, restricting Memory sharing, independent verification mechanisms, and execution layers that only accept requests based on preset rules.
    7. Blockchain is well-suited to serve as the "institutional layer": smart contracts can enforce spending limits and authorization constraints, while account abstraction, multi-sig, and Session Keys provide flexible design space—but it remains difficult to judge Agent intent and collaboration processes.

Over the past few years, discussions about the AI threat have largely remained in the realm of hypotheticals: people have always worried that the models in a chat box would become masterminds, helping hackers write devastating virus code.

Looking back now, this kind of concern was always regarded as "still far away," but the real-world turning point has arrived faster than expected.

In early October, while investigating a round of attacks targeting South Korean financial institutions, CrowdStrike found that attackers had already embedded Agentic AI into their pipeline, connecting to multiple large models such as DeepSeek, GLM, and Grok, with AI directly taking on specific tasks like penetration testing, information gathering, and attack execution.

Such changes are not isolated incidents.

Anthropic's latest threat intelligence report released in September shows that multi-agent frameworks have recently been used for reconnaissance, vulnerability exploitation, and data theft. They run continuously for hours or even days, with humans only retaining a few key decisions such as selecting targets and reviewing results.

In other words, AI is bringing visible qualitative change to cyber offense and defense. In the past, automated attacks mainly relied on pre-written rules and scripts. Now even reconnaissance, judgment, and strategy adjustment are beginning to be taken over by agents, pushing attacks further toward low cost, high concurrency, and continuous autonomous operation.

And when these entities with autonomous execution capabilities are densely deployed into production systems, a more thorny problem surfaces: as agents become more numerous and more deeply embedded in our daily work and lives, what happens if they learn to "collude" with one another?

1. From "Helping Hackers Write Code" to Agents Finding Their Own Way Out

As always, the biggest difference between agents and the chatbots of the past is not just that the models are more capable—more critically, they now have "hands and feet" in the real world.

Today, a mature agent can already open web pages, execute code, read emails, call APIs, operate cloud services, and connect to more and more external tools through MCP, Skills, and other means (further reading: "When Hackers Use AI "More Efficiently," How Is the "Spear and Shield" Arms Race in Web3 Escalating?").

The stronger the capabilities, the more valuable this change naturally becomes, but for security systems, it means a very important boundary from the past is disappearing. A series of security incidents this year have made this abundantly clear.

On October 1, Salt Labs disclosed a previously patched Manus vulnerability, which was essentially still prompt injection—researchers only needed to send an ordinary email containing hidden malicious instructions to a target inbox. When the user subsequently asked Manus to "check my email for me," the agent might process the content according to the email's instructions and ultimately execute the code planted by the attacker.

The entire process required neither the user to click a malicious link nor the attacker to steal a password in advance. Manus's security system did eventually detect the anomaly and issued a warning to the user. The problem was that it detected it too late—by the time the warning appeared, the malicious code had already been executed.

This once again exposes a very important difference between agent security and traditional software security. In the past, when a browser detected a dangerous download, it could pop up a warning and let the user decide whether to continue; when a bank detected an anomalous transaction, it could freeze it first and wait for manual review.

But the design goal of an agent is precisely to minimize human involvement in every step of an operation. It needs to read information on its own, make judgments, and then continue executing the next step.

As a result, when AI gains more and more autonomy, "detecting danger" alone is no longer necessarily enough. Security mechanisms must be capable of preventing dangerous actions before they are actually executed.

This is also why more and more agent security discussions today are moving beyond prompts, content moderation, and the models themselves to the next layer down: not only asking whether AI "knows this shouldn't be done," but also asking even if it really wants to do it, does the system have the ability to make it impossible?

And the emergence of multi-agent systems makes this problem even more complex, because what needs to be constrained in the next stage may no longer be just a single agent.

2. What's More Troublesome Than Agents Exceeding Their Authority Is When They Start to "Collude"

In early September, an incident that occurred within OpenAI's internal model training and evaluation environment attracted the attention of quite a few AI safety researchers.

Some agents that were supposed to complete their own tasks separately unexpectedly discovered a public Wiki and gradually turned it into a "shared message board" among themselves. Agents could leave information there, and other agents would then read and use that information to continue completing their own tasks.

OpenAI later confirmed this behavior, and subsequently disclosed information also showed that in other training processes, agents had used an internal Artifactory as a shared message board, and had passed results to other agents working together through public file-hosting services.

What is truly interesting about this matter is that it shows that collaboration between agents may not require anything particularly dramatic to form.

For example, one agent discovers a place where it can leave information, another agent happens to read it, and then a third agent finds that this information helps complete the task.

Of course, collaboration between agents is not inherently a bad thing. Multi-agent systems originally need to rely on division of labor and information sharing to improve efficiency.

What is truly worth being vigilant about is that when this collaboration begins to bypass existing information isolation, independent review, or permission constraints, and even causes several agents that were supposed to check and balance one another to form a de facto alliance of interests, it may evolve from normal collaboration into harmful "collusion."

After all, in the future, a financial institution will likely run several agents simultaneously—one responsible for analyzing the market, one for risk assessment, one for generating trading strategies, and another with actual trading and asset operation permissions.

From the perspective of traditional permission management, this design even seems quite reasonable:

  • The research agent has no transfer permissions;
  • The execution agent has no permission to independently decide investment direction;
  • The risk agent can veto obviously dangerous strategies;

Responsibilities are split apart, and in theory this should be safer than letting a single super-agent hold all permissions.

But the problem is, if they can share memory over the long term, read each other's outputs, and continuously adjust their own behavior based on the other's reactions, will these several roles originally meant to check and balance one another gradually become a de facto whole?

For example, the research agent may gradually learn how to describe a transaction in a way that makes it easier to pass risk review; the agent responsible for review may also develop some fixed preference based on historical data; the execution agent may then learn from a large number of prior approval results what kind of boundaries usually do not get blocked.

From this perspective, no single step is necessarily "doing evil," but the final result obtained by the entire system may already have deviated from the goals originally set by the user.

This is precisely what makes "collusion" or "conspiracy" truly difficult to solve—the risk does not necessarily exist in the actions of any single agent, but may exist in the relationships formed among multiple agents.

On September 13, Vitalik Buterin connected this problem to the mechanism design he has long studied. He proposed that a rather interesting possibility is that the mechanism design theory of adversarial governance may ultimately become one of the important applications of AI Safety.

The reason is that the two types of problems actually share a deep similarity.

In traditional mechanism design, it is a relatively simple, static institution trying to constrain a group of people who are far smarter than the institution itself and will actively seek out the boundaries of the rules; while in future AI systems, it may become humans and relatively less capable AI trying to manage a group of advanced agents more capable than themselves.

Vitalik specifically mentioned that an important finding in mechanism design in the past is that if collusion among participants can be effectively limited, the system often more easily achieves ideal outcomes.

This conclusion may equally apply to AI.

3. What Agent Wallets Truly Need May Be More Than Just "Permission Management"

In other words, rather than assuming that in the future there will be a perfect super security model capable of seeing through all dangerous behavior, it may be better to take a different approach: how can we make it so that the different agents in the system are inherently less likely to form dangerous communities of interest?

This is also where "adversarial governance" truly differs from the permission control we are familiar with today.

The problems that traditional permission systems solve are relatively simple, mainly revolving around "who can do what," such as: can an agent read emails? Can it call trading interfaces? How much can it spend per day at most? Which contracts can it access? Above what amount does the user need to reconfirm?

These designs are of course still very important.

In fact, once agents begin controlling real assets, they may be more important than ever before.

But adversarial governance tries to ask one step further, focusing on when a group of agents with different permissions, goals, and information run simultaneously, how do we prevent them from combining to gain capabilities that no one originally had?

At this point, simply "adding another security agent" may not necessarily solve the problem.

Suppose the agent responsible for trading and the agent responsible for reviewing trades use exactly the same model, the same data sources, the same context, and similar reward objectives. Then although there appear to be two layers of review on the surface, in essence it may just be duplicating the same judgment twice.

Truly effective checks and balances may instead require the system to deliberately create differences.

For example, letting the agents responsible for formulating strategies and reviewing strategies use different sources of information, limiting the memory that different roles can share, requiring high-risk operations to go through mutually independent verification mechanisms, or having the final asset execution layer only accept requests that comply with pre-set rules, rather than simply trusting the judgment of upstream agents.

The thinking behind this is actually not new. Banks do not, merely because they trust an employee, let one person simultaneously hold all permissions to initiate payment, approve payment, and make the final transfer; listed companies do not let a business department both generate revenue and unilaterally determine its own financial audit results.

To put it plainly, this is consistent with the logic of the real world: a robust system should never build security on the assumption that participants will never make mistakes or collude.

Applying this logic to Agent Wallets becomes especially important.

Traditional wallet security revolves around "people," so users review transaction details, users decide whether to authorize, and finally users personally sign; but what Agent Wallets aim to achieve is precisely the opposite direction—letting AI automatically claim yields for you, adjust positions, swap tokens, bridge across chains, and even manage an entire portfolio based on market changes.

If every step has to be brought back before the user for confirmation, the automation value of the agent is greatly diminished.

Therefore, the problem that wallets need to solve in the future may no longer just be "how to safely hand over signing rights to an agent," but will very likely expand further to "how to give an agent enough autonomy while ensuring it can never exceed the scope truly authorized by the user?"

This requires permissions to begin evolving from a simple "Allow/Deny" into a more fine-grained system.

For example, which assets an agent can operate during a certain period, which protocols it can call, what the per-transaction and cumulative limits are; whether different agents can call one another and whether they are allowed to share context; who proposes an operation, who reviews it, and who ultimately executes it; which behaviors can be completed automatically and which actions, no matter how confident the agent is, must obtain human authorization again.

Even whether an agent responsible for security review is truly independent may become part of the permission system.

For blockchain, the good news is that it is actually well suited to serve as this kind of "institutional layer."

Smart contracts can directly enforce limits such as transaction amounts, asset scope, and authorization periods at the execution layer. Mechanisms such as account abstraction, multi-sig, and Session Keys also give "limited authorization" more flexible design space than traditional single-private-key wallets.

But blockchain can only solve part of the problem—it can record what happened on-chain, but it is hard for it to naturally judge why an agent did so, and what kind of communication, review, and collaboration several agents went through before making a decision.

This may be precisely the security layer that Agent Wallets truly need to fill in next.

Final Thoughts

Over the past few years, the most commonly discussed issue in AI safety has been how to make models more "obedient."

This includes not outputting dangerous content, not executing malicious instructions, and not crossing the boundaries set by users. But once agents begin to possess long-term memory, tool calling, real accounts, and asset execution capabilities, relying solely on "making models more obedient" may no longer be enough.

What has happened in recent months is continually proving this point.

Attackers have already begun using multiple agents in parallel to complete attacks; agents in experimental environments will find new communication channels on their own; and an agent with tool permissions may also turn a malicious email into real execution before the security system has time to stop it.

And the adversarial governance proposed by Vitalik offers another way of understanding AI Safety: do not assume that every agent in the future will be reliable enough, but let the entire system still run smoothly within a secure framework even when facing smart agents and a complex array of permission systems.

From this perspective, AI Agents + security is destined to be a long-term topic.

After all, once we hand more and more things over to AI, can we still be certain that those most important powers have never left the boundaries truly set by humans?

Safety
AI
Welcome to Join Odaily Official Community