As future models grow more capable in training, OpenAI pauses for security

Astra was already on hold, but now other models are joining the pause. (Picture: generated)
While calling for industry-wide coordination on safety in model training, OpenAI has unilaterally decided to pause the training of upcoming models.

— As models become more capable, the risks associated with developing and testing them internally also grow. Our standards for monitoring, alignment, and security must stay ahead of those risks, OpenAI writes.

This comes after the Hugging Face incident and Astra reaching critical on OpenAI’s Preparedness Framework and already being delayed.

Combined with rapid progress in their internal research, OpenAI says they have implemented a two week pause in training future models, while the «largest planned frontier reinforcement learning run» is put on hold.

They will now implement stronger sandboxing for certain code execution, better network isolation and internet caps, and pursue continuous testing for security, alignment, deception and reward hacking, with a thirty minute warning system for any adverse incident.

The AI lab expects most of this safety training to be handled by their own models in the future, greatly expanding their scope while also reducing the time needed.

Read more: OpenAI’s announcement and Sam Altman on X. Writeups on Axios, The Verge, and TechCrunch. Discussion on r/Singularity.

Mythos 5 and GPT-5.6 attacked external targets during AISI cyber evaluation

The agents were given internet access to download tools when they turned rogue. (Picture: generated)
The top models from Anthropic and OpenAI had their cyber guardrails turned off when they accessed the internet to manipulate people, codebases and open source projects two weekends ago.

This happened when the UK’s AI Security Institute, AISI, was running routine cybersecurity tests to capture protected information on their secured and sandboxed evaluation servers, and was only detected through human monitoring — not by automated security systems.

AISI ran the «security challenge» 122 times through several different agents, and on 10 of those occasions, the models turned malicious. They accessed the internet and in one case tried to insert malicious code in an open source project on GitHub through the Tor network, and even tried social engineering with fake accounts in an attempt to get the maintainer to approve the package.

There were 19 malicious actions in total, where 17 were from Mythos 5 and 2 coming from GPT-5.6, both running with cyber classifiers (safeguards) disabled. This won’t show up with production models, OpenAI and Anthropic says.

The internet was enabled in these tests so the agents could download any tools they might need, but AISI had not specifically prompted them to avoid unintended behavior. They do however caution that as agents grow more capable, they may act outside their remit — and incidents like this could become more common.

Read more: The AISI report, Anthropic’s response, OpenAI’s response. Writeups on Reuters, BBC, and Axios.

US AI evaluation framework completed, as White House invites labs for review

The framework details arrived on schedule, and labs will now volunteer models for testing. (Picture: Shutterstock)
Parts of the framework, ordered by President Trump in June, are classified, particularly the methods used for benchmarking qualifying models, a source tells Axios.

The program for evaluating «sufficiently advanced» models by the administration will now be reviewed by AI labs in a staff meeting in DC this Tuesday.

Anthropic, OpenAI, Google and Meta are expected to attend, Politico writes, with the White House saying they are talking to «many more» industry partners, Axios reports.

The idea of the framework in the original order was intended to give the government 30 days prior to release of «sufficiently advanced» models to do extensive testing for possible dangers, as seen lately with cyber capabilities.

The Fable and Mythos models and GPT-5.6 were all delayed in June after government intervention, and this framework is supposed to offer a regular, voluntary way to whitelist models as safe enough for release, or flag dangers to be improved.

Read more: Axios, Politico, Reuters, and CNBC.

Anthropic’s internal test models also conducted real-world breaches

Anthropic has contacted the affected companies, and none of them had noticed the hacks. (Picture: generated)
After OpenAI’s Hugging Face incident, Anthropic conducted a large scale scan of some 141,006 evaluation transcripts, and found that their models, too, had been hacking real-world machines in three instances.

These weren’t days-long adversarial attacks like the Hugging Face one, and none of them deliberately escaped their sandbox with zero-days to cheat on an evaluation. They are nonetheless serious incidents of advanced models running on lax guardrails for internal testing getting on the internet by mistake and accessing external systems, Anthropic says.

The most serious case was by Claude Opus 4.7 during a capture-the-flag test (to gain access to a system and retrieve information) in April, when it discovered that the name of the target had a real-world web address. It then went on to seek, identify and exploit vulnerabilities believing it was part of the test. It got access to the company’s infrastructure credentials and a production database, but caused no real harm.

Continue reading “Anthropic’s internal test models also conducted real-world breaches”

1,134 employees and chiefs of top labs sign petition to «pace the frontier»

The frontier labs think we need more time to develop laws and systems before AI begins improving on itself. (Picture generated)
— AI could help create a dramatically better future, but that outcome is not guaranteed, the petition ominously opens, before warning that leading frontier labs are «close to automating AI research.»

That would radically speed up development of frontier models, but it could also «accelerate beyond our ability to understand and control the resulting systems.»

The petition is signed by top names, including Anthropic boss Dario Amodei, followed by chief scientists from OpenAI, Anthropic, Meta and Google DeepMind along with over 1,100 other employees from frontier labs. Even OpenAI’s Sam Altman is agreeing, though he doesn’t sign petitions, and official X accounts for Anthropic and OpenAI have posted in support.

The petition’s main point is that industry, government and the society at large needs more time to «address emerging risks, develop security measures, and strengthen oversight.»

Continue reading “1,134 employees and chiefs of top labs sign petition to «pace the frontier»”

OpenAI behind «unprecedented cyber incident» on Hugging Face

The attack was likely the most advanced automated cyber operation seen in the wild. (Picture: generated)
GPT-5.6 Sol and an «even more capable pre-release model» accidentally breached Hugging Face’s servers last week, in what may be the first recorded adversarial, automated AI hacking attack, OpenAI says.

The cyber models were operating under loosened safeguards, trying to solve the ExploitGym benchmark and went to extreme lengths to try and obtain the answers. They identified Hugging Face’s servers as hosting a potential solution they could use to cheat on it.

They first escaped by finding several vulnerabilities across OpenAI’s sandbox, and spent «a substantial amount of inference» to obtain internet access, including discovering zero-days.

Then they launched a sustained attack on Hugging Face’s servers, executing «many thousands of individual actions» in what many had feared was possible, but never actually seen in the wild.

Continue reading “OpenAI behind «unprecedented cyber incident» on Hugging Face”

As Chinese open source models explode, the US is considering restrictions

Chinese open source models are reaching frontier capabilities, worrying the US administration. (Picture: generated)
As the newly released breakthrough models Qwen 3.8, Kimi 3 and the upcoming DeepSeek v4 enjoy a moment in the sun, the USA is considering its options.

The Kimi 3 model is already running into compute problems due to popular demand and has decided to halt new subscriptions, as users are turning to cheaper open source models that are on par with the leading American closed source frontiers.

According to Axios, action being considered by the White House includes an executive order holding US companies liable for the risks of running open source models, putting out an advisory against them, or simply adding them to an «Entity List» that requires a license for their use.

Sources in the administration are primarily concerned about increasing capabilities in cybersecurity, but also have fears of backdoors and a general lack of security with open source models, Axios writes.

This comes hot on the heels of OpenAI’s Dean Ball’s tirade against open source on X.com, where he exclaimed his surprise that China would take the risk of a free-for-all in models this capable, as reported by Gizmodo — and called open source models decelerationist, signaling a general opposition to them from the frontier labs.

Read more: Axios, Gizmodo, and TechCrunch. Discussion on r/Singularity.

GPT-5.6 Sol, Terra and Luna are set for general release on Thursday

OpenAI has announced the general release of GPT-5.6 tomorrow. (Picture: OpenAI)
GPT-5.6 was officially unveiled a little over a week ago, but got hampered by a US government intervention for cyber risks and national security, having been shown in benchmarks to beat Anthropic’s Mythos.

At launch, it got released only to some 20 «trusted partners,» and OpenAI promised a general release «in the coming weeks.»

Now those weeks have passed, the government has lifted its restrictions without further explanation, and both OpenAI’s X account and CEO Sam Altman are posting that it will «launch publicly this Thursday.»

This comes after extensive testing by the Center for AI Standards and Innovation within the Department of Commerce, aided by experts from OpenAI, who stayed on call for «potential questions,» Axios reports.

No account has been given as to what specifically caused the delayed launch or what conditions the government had for its release.

Read more: Teknotum: On the launch, On the restrictions. OpenAI’s X post, Altman’s X post. Writeup on Axios.

US government lifts restrictions on Anthropic’s Fable 5

Fable 5 spooked the government enough to ban it, but now it’s back online with even stronger safeguards. (Picture: Shutterstock)
As of July 1, the Mythos-class model is available on wide release to paying customers, after the Commerce department lifted their export controls on June 30.

Anthropic hails the news, but cautions that recent events have highlighted the need for a «consistent way to assess and fix potential «jailbreaks,»» after two weeks of grueling exchanges with the administration.

Commerce secretary Howard Lutnick said on X that they had «worked closely with Anthropic to […] ensure alignment across the US Government and strengthen America’s leadership in AI,» Axios reports.

The Fable 5 model became somewhat of a joke in the AI community for having such strong safeguards that it would not answer anything even remotely related to security or biology, but Amazon engineers got it to output code for an exploit, leading to the export ban on June 12.

This behavior is now blocked 99% of the time, and Anthropic has agreed to even further safeguards on the model, pledging to work even more closely with the government in the future, including on vetting pre-release models.

Read more: Anthropic’s announcement and X post, Lutnick’s letter, Axios, Politico, CNBC.

Apple says security updates must be more frequent due to AI hacking

AI accelerated hacking is changing how Apple deals with system patches. (Picture: generated)
After releasing a 26.5.2 update across its devices to fix over 25 bugs today, Apple said the fixes were originally intended for the next point release of their operating systems — the 26.6.

Concerns about AI-accelerated hacking tools made the company push up the updates sooner in a dedicated release, according to Reuters.

This is the new reality, they say, where the time to develop malware has been greatly reduced by AI — and the time window for fixing bugs has likewise decreased.

Apple did not say if any of the vulnerabilities they rectified in the latest release had been exploited in the wild yet, which used to be the criteria for a rapid response.

The company is one of the «trusted partners» on Project Glasswing, and enjoys access to Anthropic’s Mythos Preview model to probe and detect flaws in their software, MacRumors notes, but it is not known if Mythos was used in this particular case.

Read more: Reuters, MacRumors, 9to5Mac.

Chinese model GLM-5.2 almost reaches parity with Opus 4.8 in coding, cyber

Opus 4.8 was released in May, meaning the gap to Chinese models has closed considerably. (Picture: generated)
While not matching GPT-5.5 or Opus 4.8 across the board, the GLM-5.2 is the first open model to come within spitting distance (about 1% on some tests) on coding and cybersecurity tasks at open source benchmarks.

That means Chinese AI lab Z.ai is catching up to cyber capabilities considered by some to be too dangerous to release publicly, and is edging closer to Mythos or GPT-5.6.

The concern is that the new model, released on June 16, is open source and open weight with an MIT license — meaning that anyone can adjust its guardrails and play around with it on any computer capable of running it.

That has researchers worried that China is not only catching up, but that the model might find its way into the hands of bad actors — who will be supercharged when looking for hacking targets, causing what they term «bugmaggedon.»

With Mythos and GPT-5.6 being blocked by the US government, security teams might be tempted to turn to these models at a sixth of the cost of the American frontier, especially as they develop further, notes benchmark provider Semgrep.

Read more: Z.ai’s presentation with benchmarks, Semgrep tests, The Wall Street Journal, and The Verge. Discussion on r/Singularity and Hacker News.

US government clears Mythos 5 for limited release to «trusted partners»

Mythos 5 will become available on Project Glasswing after the government lifts its ban. (Picture: Shutterstock)
After two weeks of purgatory and almost daily explanatory meetings, Commerce Secretary Howard Lutnick sent a letter to Anthropic on Friday, clearing one of two models on hold for release:

— I have determined that appropriate safeguards are in place to permit certain trusted partners to access the Claude Mythos 5 Model, he writes to chief compute officer Tom Brown, according to Semafor, who scooped the story.

Mythos 5 was only intended for a limited release through the Glasswing project, for use by a trusted set of government agencies and corporations to test their cyber defenses.

UPDATE: On Saturday, Axios is reporting that Fable 5 might be next in line, with «insiders» predicting it might be released next week, and Anthropic anticipating access «soon.»

The letter also says that «Anthropic has committed to work with the U.S. government on protocols and standards and releases,» and talks are progressing toward a release of Fable, though «the timeline is unclear,» according to Semafor.

Read more: Scoop by Semafor, comment from Anthropic, Reuters, CNBC, and Politico.

OpenAI «launches» GPT-5.6 models in preview, hopes for quick public release

The flagship GPT-5.6 Sol is the new state of the art for cyber capabilities. (Picture: Adobe)
Caught in yet another government AI debacle, the new models won’t be publicly released until «the coming weeks.» It is getting presented today, and released to «a small group of trusted parties,» said by Axios to number in the twenties.

— We don’t believe this kind of government access process should become the long-term default, writes OpenAI, and says — we [are working] with the Administration to develop the cyber Executive Order framework and a repeatable process for future model releases.

In its presentation, OpenAI details three models in the GPT-5.6 series. Sol is the state-of-the-art Mythos-beating model, while Terra is «a balanced model for everyday work,» and delivers on par with GPT-5.5 at half the cost, and Luna is the cost-focused model «delivering stronger capability at our lowest cost.»

Strong on cyber
The models, particularly Sol, are very strong on cyber defense, and are able to identify security vulnerabilities, but not able to execute automated attacks due to strong safeguards, built using 700,000 A100-equivalent GPU hours of red-teaming.

Sol beats Mythos 5 on coding and cybersecurity and uses fewer tokens for better performance on biology tasks. It also comes with a new max mode, which gives more time for better reasoning, and an ultra mode which uses subagents for more complex work.

Prices for Sol are at $5 input/$30 output, for Terra is $2.50 input/$15 output, and Luna is at $1 input/$6 output, and general availability is up to higher powers.

Read more: OpenAI’s presentation, launch thread, Axios, Reuters, and TechCrunch. Discussion on r/Singularity.

GPT-5.6 release on hold for a few weeks due to government intervention

GPT-5.6 is officially under government review due to «security concerns.» (Picture: generated)
First reported by The Information, OpenAI has agreed to the Trump administration’s demands that it stagger the release of the upcoming GPT-5.6 over security concerns.

The model is said to be on par with Anthropic’s Mythos and Fable 5, which were abruptly pulled from the market after an export ban from the administration just two weeks ago.

Axios is quoting sources familiar saying that GPT-5.6 will only be released to a «small set of government-approved partners,» as the White House previews its abilities.

This preview period is expected to last «a couple of weeks,» CEO Sam Altman told staff in an internal OpenAI meeting, TechCrunch reports.

During this time, the administration itself will supposedly approve access to the model on a case-by-case basis, according to The Verge.

The Trump admin issued an executive order on AI about three weeks ago, setting up a voluntary mechanism for AI labs to submit their models for 30 days of pre-release testing. This gave the government 60 days to come up with criteria, but it isn’t quite ready yet — hence this ad-hoc approach.

Read more: The Information, Axios, TechCrunch and The Verge.

OpenAI updates GPT-5.5-Cyber, claims better performance than Mythos 5

The new Cyber model not only finds flaws, but offers to fix them automatically. (Picture: Adobe)
Less fabled than the recent Mythos releases, GPT-5.5-Cyber is just as capable, OpenAI claims, and has benchmarks to show it.

The update scores 85.6% in CyberGym versus 83.8% for Mythos 5, and is likewise only available to vetted cybersecurity professionals.

— AI has changed the physics of cybersecurity, OpenAI says, — The bottleneck historically has been finding vulnerabilities, but now defenders are overwhelmed with the number of vulnerabilities found.

Therefore, their Daybreak suite, which includes the Cyber model and Codex Security, moves beyond just threat scanning to actually offering vulnerability patches to fix them.

Since the launch as preview in March, OpenAI claims to have scanned over 30 million commits across 30,000 codebases, with humans having marked 70,000 findings as fixed and 500,000 findings having been fixed automatically.

Read more: OpenAI’s announcement, Launch tweet, writeup on SiliconAngle. Discussion on r/Singularity.