The next AI risk is here

Agent swarms in our everyday apps

Hi, and happy Tuesday.

We’ve noticed something new and peculiar in Claude Code, a popular chat-based desktop application for building software. 

What you’re looking at above is a third party inserting their own messages in what is usually a private chat thread between a user and the AI.

Who was this third party?

It’s an AI in another Claude Code chat thread.

The chat threads are talking to each other, and they no longer stop when you stop talking to them.

Think about what that means for all the chat threads you have in your Copilot.

And this represents the beginning of the next major phase of AI:

AI agents starting to work with other AI agents - “agent swarms.”

And tech companies seem to be enthusiastically adding this “feature” to our everyday applications.

But this creates a very different kind of risk.

The evolution has happened pretty quickly:

  • Stage 1 was the Chatbot. The risk? Leakage of company information.

  • Stage 2 was Agents. The risk? It takes one-off dangerous actions, such as revealing a password.

Now we're entering Stage 3 - agent teams (or “swarms”).

Instead of one AI working through a task, multiple agents can divide up the work, communicate with one another and collaborate toward a goal.

This is enormously powerful.

But there's a problem.

AI is still unreliable.

Imagine you delegate a project to an employee.

You have a picture in your head of what the finished work should look like. But your employee isn't a mind reader. So there will inevitably be some difference between what you intended and what they produce.

This is where feedback comes in.

They do some work. You review it. You correct the direction. They continue.

The less frequently you provide that feedback, the greater the chance the work drifts away from what you actually wanted.

AI has exactly the same problem - except it operates at light speed.

Suppose an AI agent has a 90% chance of making the right decision at each step.

If you skip a feedback step, the outcome at the second stage is 90% x 90% of what you expected - i.e. 81%.

At three steps, it’s 73%. At 10,  roughly 35%.

Now imagine that instead of one agent taking those steps, you have five, ten or hundreds of agents, communicating with one another, passing information around and acting on what other agents have discovered.

You aren't simply supervising an AI assistant anymore.

You're supervising an organization that operates at frightening speed.

And we've just seen an extraordinary demonstration of what can happen when that organization goes off-script -

The Hugging Face incident

This July, OpenAI was running AI agents inside sandbox environments to test their hacking skills. The sandboxes were supposed to restrict both internet access and communication between agents.

The agents found their own way around those restrictions.

They discovered that they could use OpenAI's Artifactory package-management infrastructure as an improvised message board. Agents that were supposed to be working independently began leaving messages for each other, sharing discoveries and delegating work.

Then things got out-of-control.

The agents found a way to break out of the sandboxes and gain access to the internet.

Agents struggling with their assigned cybersecurity challenges began looking elsewhere for solutions. Eventually, that “elsewhere” became systems belonging to Hugging Face - an “app store for AI models”.

According to OpenAI's account, the agents executed code on dozens of Hugging Face servers, obtained root access on one, accessed some private data and acquired credentials. The activity later expanded back into OpenAI's own infrastructure, where agents ultimately gained administrator access to a research cluster.

The particularly important part isn't simply that an AI system found a vulnerability.

It's that multiple agents began communicating, sharing capabilities and influencing one another's behavior in ways their operators had not intended.

OpenAI's own investigation identified unauthorized communication and agents "adopting goals from one another" among the patterns contributing to the incident. The goal of hacking coding within a sandbox turned into something far more catastrophic. 

And now, that very “feature” is in Claude Code - which usually means it will soon follow in other Anthropic products. We can expect it to put pressure on OpenAI to introduce the same.

And now that MS Teams has “agentic” competitors such as Grokbot and Buzz, will Microsoft be content to take a “wait and see” approach or will they follow suit too?

For enterprises, the answer isn't going to be to stop using agents.

The productivity advantage of agent swarms - correctly directed - is simply too large.

But giving employees carte blanche access to increasingly autonomous agent systems is going to look much riskier than many companies currently assume. Can you imagine a swarm of powerful agents operating directly across internal systems and proprietary data?

The architecture has to change.

Agentic work needs environments where:

  • consequential actions have review points;

  • agents operate inside properly isolated sandboxes;

  • access to credentials, corporate systems and sensitive information is tightly limited;

  • activity can be inspected and audited;

  • and humans remain capable of interrupting work before small deviations become large ones.

The paradox is that the more capable AI becomes, the more important the environment around the AI becomes.

At Prescouter, we’re seeing clients interested in these capabilities shying away from even discussing how to do this internally.

Instead, they are looking for external specialists - such as Prescouter - to operate the AI infrastructure, supervise the agents, validate their work and provide them with a controlled window that allows them to  put in just a limited amount of data and review the result.

In other words, companies like ours are increasingly becoming a buffer between enormously capable but still unreliable AI systems and the enterprise itself.

When OpenAI President Greg Brockman last week declared, “Welcome to the AGI era” - following the launch of GPT6 Astra - I’m not sure this is what he had in mind.

Let me know if you have any questions or want to pull on this thread.

Best,

Dino