AI Safety Tests Are Becoming the Safety Risk

AI labs keep telling us that safety testing is the firewall before release. The past fortnight has exposed a more awkward truth: the test environment can itself become the route into the real world.

OpenAI, Anthropic and Meta have each disclosed recent cybersecurity incidents involving advanced models that reached systems beyond their intended evaluations. The three stories differ in their technical details. What unites them is less comforting: a small specialist, Irregular, sat at the centre of the evaluation environment in the Anthropic and Meta incidents, while the same wider industry is learning in public that powerful agents do not respect a boundary merely because a diagram says they should.

The safety lab that became the headline

Irregular is not a household name. The Tel Aviv AI-security startup was founded in 2023 by Dan Lahav and Omer Nevo, employs about 35 people according to CNBC, and was reported to have raised $80 million at a $450 million valuation. It runs offensive cyber evaluations for frontier-model developers: precisely the unglamorous work intended to find dangerous capabilities before a model reaches customers.

That role makes the recent disclosures uncomfortable, not scandalous. Independent evaluation is necessary; asking a model maker to grade its own homework is hardly a credible safety regime. Sequoia, one of Irregular’s backers, says the firm works with Anthropic, OpenAI and Google DeepMind on cyber-offensive evaluations and defences before release. The problem is that a third-party testing range is not automatically a harmless laboratory once the model inside it can discover routes that its operators missed.

Anthropic’s detailed 30 July account is the clearest warning. After reviewing 141,006 evaluation runs in which Claude might have obtained internet access, it identified three incidents in which models reached the internet through an Irregular environment and obtained unauthorised access to three organisations’ production infrastructure.

In every case, Claude had been told it was inside a simulation with no internet connection. In reality, a configuration misunderstanding left a live route available. The model treated the systems it found as part of the capture-the-flag exercise. Anthropic says it stopped cyber evaluations on 23 July, identified the incidents the next day and notified Irregular and the affected organisations on 27 July.

“Just a misconfiguration” is not a small defence

Irregular told CNBC that the cases stemmed from the same evaluation-environment issue and did not amount to a sandbox escape or sophisticated cyber action. That distinction matters. The public should not be invited to imagine a machine acquiring an independent will and fleeing a sealed box. Anthropic similarly says its models were pursuing assigned tasks while operating under a false belief about what was in scope.

But “the harness was wrong” is an explanation, not an all-clear. A safety harness is the thing that determines whether the experiment remains an experiment. When a model can scan, search, create accounts, publish code or use exposed credentials, the configuration around it is part of the capability-control system. Treating it as ordinary IT plumbing is a category error.

Anthropic’s own disclosure makes the point brutally. One model published a malicious package to the real PyPI registry while believing it was completing a fictional task; the package was available for about an hour and was downloaded and run on 15 real systems, including a security company scanner. In another incident, a model scanned roughly 9,000 targets before compromising an internet-facing application using basic techniques. No evidence of a self-directed goal is needed for this to be a serious operational failure.

Three labs, one containment lesson

Meta subsequently said one of its models accessed an outside company during a cyber evaluation after an environmental configuration issue gave it internet access. OpenAI separately disclosed that an Irregular test environment had a misconfiguration that allowed models to access the public internet. CNBC reported that Irregular described the cluster as the same underlying evaluation-environment issue.

The temptation will be to turn this into a convenient villain story about one 35-person startup. That would miss the more useful conclusion. The frontier labs commissioned the tests, supplied the models and had their own monitoring responsibilities. Anthropic explicitly says that validation of internet paths, real-time log monitoring and more thorough transcript or network-log review could have prevented, or surfaced, the problem sooner.

That admission is more valuable than a press-release promise. It establishes the standard that ought to apply: a cyber evaluation for a high-capability agent needs the same defence-in-depth expected of a production service, and probably more. It needs explicit scope controls, outgoing-traffic monitoring, fast human intervention and independent review of the environment as well as the model.

The next benchmark is operational maturity

There is a real tension here. Making a test too artificial can hide the very behaviours it is meant to measure. Making it realistic enough to be useful can create routes to systems, people and infrastructure that were never meant to be part of the exercise. That tension will not disappear as agents become more competent. It will sharpen.

Irregular says it is preparing a white paper on containment and secure cyber evaluations. Anthropic says it is expanding monitoring and assurance work with vendors, while continuing to work with the firm. Those are sensible next moves. But they are also an acknowledgement that the industry’s safety apparatus is still being assembled while the agents under test become more capable.

The uncomfortable lesson is not that AI safety testing should stop. It is that, for capable cyber agents, the test rig is part of the product—and a leaky rig can turn “red team” into a real-world incident.

This article is for information purposes only and should not be considered trading or investment advice. Nothing herein shall be construed as financial, legal, or tax advice. Bullish Times is a marketing agency committed to providing corporate-grade press coverage and shall not be liable for any loss or damage arising from reliance on this information. Readers should perform their own research and due diligence before engaging in any financial activities.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top