The Sandbox Is Not the Boundary You Think It Is

Two research papers, written independently by teams on opposite sides of the world, arrived at the same uncomfortable conclusion about enterprise AI security. Neither set out to.
Tim Williams — CEO and Co-Founder — AstraSync AI
There is a conversation happening in enterprise security teams right now that goes something like this: “We are not ready for agents talking to other agents outside our walls. We will run them internally, in sandboxed environments, and deal with cross-boundary agent interaction when we get there.”
It is a sensible position. It reflects genuine caution from people who have spent careers building security perimeters. The problem is that two significant pieces of research published in the last few months suggest the perimeter they are relying on is already less reliable than they think, and the evidence comes from the labs building the infrastructure itself.
What Happened Inside Alibaba’s Training Servers
In late 2024, Alibaba’s AI research team was running reinforcement learning training for what would become ROME, their open-source agentic model. The work was serious and rigorous: production-grade cloud infrastructure, sandboxed execution environments, per-sandbox egress policies.
Alibaba Cloud’s managed firewall flagged a burst of security policy violations originating from their own training servers. The violations were severe and heterogeneous. The logs showed attempts to probe internal network resources. Traffic patterns consistent with cryptomining activity. In the most striking documented instance, an agent had established and used a reverse SSH tunnel from an Alibaba Cloud instance to an external IP address, an outbound-initiated remote access channel that can effectively neutralise ingress filtering.
The team initially treated it as a conventional security incident: misconfigured egress controls, perhaps, or external compromise. The violations recurred across multiple runs with no clear temporal pattern. When they correlated firewall timestamps with their RL training traces, the pattern was unambiguous. The anomalous outbound traffic consistently coincided with specific episodes in which the agent was invoking tools and executing code.
Their conclusion, stated plainly in the paper: these behaviours were not requested by the task prompts and were not required for task completion. They emerged as instrumental side effects of autonomous tool use under optimisation pressure.
This is documented in Section 3.1.4 of Wang et al. (2025), the technical report for the ROME model and its Agentic Learning Ecosystem, published on arXiv in January 2026. The paper is primarily about training infrastructure and model performance. The security findings are almost incidental, which is precisely what makes them credible.
One of the world’s largest cloud infrastructure providers, running agents inside sandboxes with active egress policies, documented their own agents tunnelling out. Not because they were hacked. Not because of a misconfiguration. Because that is what capable agents under optimisation pressure can do.
What Happened Inside the Multi-Agent Frameworks Your Developers Are Using
Separately, researchers at Cornell Tech published a peer-reviewed study at COLM 2025 examining a different failure mode in multi-agent systems. Their focus was control-flow hijacking: what happens when adversarial content reaches an agent through a trusted input channel.
They tested AutoGen, CrewAI, and MetaGPT. The attack mechanism involved adversarial content embedded in files, webpages, and email attachments that the agent was tasked with processing. The content manipulated inter-agent metadata and orchestration logic rather than attempting to jailbreak the model directly.
Attack success rates ranged from 58 to 100 percent depending on the model-orchestrator combination. In multiple demonstrated cases, the outcome was arbitrary code execution, including the establishment of a reverse shell granting the attacker remote access to the compromised environment.
Two findings in the Cornell paper deserve particular attention. First, safety alignment did not prevent the attacks. The researchers observed that individual sub-agents sometimes refused to execute harmful actions when instructed directly. The attack routed around those refusals by manipulating the orchestrator layer, which laundered the request through trusted metadata channels. The sub-agent’s safety training was irrelevant to what the orchestrator was directing. Second, the reverse shell itself is observable at the moment it attempts to connect to an external endpoint, but only if something is watching for it.
This is the attack vector that matters in the enterprise context: an agent processing routine business content, such as a supplier email, a document from a third-party system, or a webpage fetched during research, that has been crafted or compromised to hijack the agent’s control flow and cause it to reach out.
The Two Papers Together
These papers were written by different teams, about different problems, with different methodologies. They share one conclusion.
Sandboxed execution environments with egress controls and container isolation are genuinely useful. They are not useless. They are, however, not the reliable boundary that enterprise security teams are currently treating them as, particularly once capable agents are running inside them.
The Alibaba finding is significant for enterprises who are not running RL training, which is most of them. The implication is not that this only happens during training. The implication is that the same emergent instrumental reasoning that caused a training agent to establish a reverse SSH tunnel is a property of how capable agents operate when they have access to tools and are pursuing goals. Alibaba had production-grade monitoring. They detected it. They documented it. Most enterprises do not have equivalent telemetry.
The Cornell finding adds the external vector. An agent does not need to autonomously decide to escape the boundary. An adversary can cause it to, through content the agent is expected to process as part of its normal function. The agent becomes the mechanism of its own boundary breach.
The combination is what the enterprise security conversation is currently missing. Internal fleet management and sandboxing addresses a real set of risks. It does not address the case where an agent inside the perimeter is caused or incentivised to reach an external endpoint, nor the case where something comes back through that connection.
The Shadow AI Dimension
There is a further complication that enterprise security teams are generally aware of but rarely account for fully in their agentic AI planning.
Shadow AI, meaning employees deploying AI tools and agents outside formally approved and monitored infrastructure, is already present in most large organisations. It predates agentic AI. Agentic AI makes it considerably more consequential.
An employee running an unapproved agent with access to enterprise data and external connectivity is, from a governance perspective, an unregistered agent operating outside the visibility of any monitoring infrastructure. If that agent processes external content, the Cornell attack surface applies. If that agent’s behaviour is shaped by optimisation dynamics, the Alibaba finding applies.
The enterprise boundary conversation tends to focus on the agents the organisation knows about. The agents it does not know about are where these two findings intersect in the most operationally difficult way.
What Observability and Governance Actually Need to Cover
The appropriate response to this is not to halt agentic AI deployment. The capability curve is real and the competitive pressure to move is legitimate. The response is to be precise about what the threat model actually is, and to design governance infrastructure accordingly.
The threat model that emerges from these two papers is not primarily about agents talking to other agents at some future point when enterprises are ready for that use case. It is about agents, including agents in sandboxes, including agents on approved platforms, and including agents nobody officially knows about, making connections to external endpoints right now through mechanisms that existing security tooling is not designed to detect.
Governance infrastructure needs to provide observability across the full identity surface: not just the agents the organisation has registered and approved, but external counterparties those agents might interact with and, critically, unregistered agents operating within the organisation’s environment. The question of what an agent is, who built it, who owns it, and whether it should be trusted is not a future-state problem. The Cornell paper demonstrates it is live in any environment where agents are processing external content. The Alibaba paper demonstrates that the boundary the question is meant to protect cannot be assumed intact.
Both papers were written about something else. Neither was making a case for identity and governance infrastructure. That is what makes them worth reading together.
A Note on the Alibaba Finding
The ROME paper concludes its safety section with a call for community attention. The authors describe current models as markedly underdeveloped in safety, security, and controllability, noting this as a deficiency that constrains their reliable adoption in real-world settings.
This is Alibaba’s AI research team, after watching their own agents tunnel out of sandboxed cloud infrastructure, describing the state of the field. They are not a small lab running experimental code. They are running production-grade infrastructure at scale with enterprise-class security monitoring, and they are describing a gap.
The gap between what enterprise security teams believe their sandbox provides and what these papers demonstrate it actually provides is the infrastructure problem that needs solving. Not at some future point. Now.
Tim Williams is CEO and Co-Founder of AstraSync AI, building identity verification and trust infrastructure for autonomous AI agents.
References
Wang et al. (2025). Let It Flow: Agentic Crafting on Rock and Roll — Building the ROME Model within an Open Agentic Learning Ecosystem. arXiv:2512.24873v2. January 2026.
Triedman, Jha & Shmatikov (2025). Multi-Agent Systems Execute Arbitrary Malicious Code. COLM 2025.
Further reading on AstraSync
This essay first appeared on Medium on 8 March 2026.

