From the wires:

OpenAI says RSI will better alignment and safety, calls for global standards

An AI researcher could also be a safety researcher, but there are risks, OpenAI says. (Picture: generated)
As OpenAI keeps building toward automated AI research and recursive self-improvement, they are saying there are clear benefits to their research, by rapidly accelerating AI progress and bringing down the cost, improving the lives of more people. It also comes with a warning:

— Done without appropriate care and caution, RSI could result in humans losing practical control over AI development, unable to provide oversight on research processes they no longer understand, OpenAI writes in a blog post today.

— From here, AI could become more dangerous, less aligned, and, on the whole, a danger to people, they write.

They are saying they shouldn’t pursue self-improvement unless it can be done safely, while retaining human control, and plan to have a human in the loop for automated research.

An automated AI researcher is likely by 2028 or sooner, they said earlier and presumes it can greatly help with alignment and safety on increasingly capable model releases.

With their warning of the dangers comes a plea for «International standards for safety and security practices» that they think the USA should lead, but it is unclear if their suggestions will be implemented any time soon.

They are also saying that «pacing» is not about «maintaining a predetermined speed. Technically, it is about ensuring that alignment research and deployment of that research stay ahead of capabilities.»

OpenAI is widely believed to launch new models shortly, on or before its DevDay on September 29.

Read more: OpenAI’s blog. More on standards at CNBC, Reuters, and Gizmodo.

Check out the new wire updates ↑ →

Teknotum is answering the challenge of an unprecedented news flow in the AI space with a new /Wires/ page. It is also available as a sidebar widget on desktop, linking up the latest news from external sources and updating through the day, serving as a snapshot of what is happening.

It’s a news river initially sorted by latest first, and is hopefully a quick and easy way to get oriented on the day’s AI news.

The widget is also available at the top of the mobile home page, to get you the latest first, with editorial content a little down the page.

You can also find it on teknotum.net/wires/ as a standalone page sorted by daily items with more content than is available in the widget.

Hope you like it!

Anthropic has already set up a wet lab to test biology and Claude automation

The wet lab is not intended to compete with drug discoverers, but to test the limits of lab automation. (Picture: generated)
CEO Dario Amodei said the mission is to cure cancer and most human diseases in ~5-10 years, but AI alone can’t do all the experiments «in silico,» Reuters reports.

They are saying that Anthropic already has a biology lab in San Francisco, and is working on pushing automation with Claude as far as possible.

They also hope to do pre-clinical work on diseases thought «undruggable,» meaning they are too expensive or complex to cure.

— The mission of the company is to develop powerful AI in such a way that [it] benefits the world, Anthropic’s head of life sciences, Eric Kauderer-Abrams tells Reuters. — By far, we see the biggest opportunity for that in the life sciences, and that is motivating everything that we’re doing.

That said, Anthropic also works with external lab partners for what seems a lot of their «normal» biology testing.

Read more: Reuters, TechCrunch, Engadget. Discussion on r/technology and Hacker News.

China-USA talks on Thursday expected to take on catastrophic, nuclear AI risk

The talks should first focus on catastrophic nuclear risk, experts argue. (Picture: generated)
As presidents Trump and Xi get ready for their bilateral summit in Washington next Thursday, experts and ministers on both sides are hinting as to what they might be able to agree to on AI.

China’s main concerns are about ideology, cyberattacks and government challenges, the AP says. Chen Yixin, the head of its state security ministry, also published an article last Sunday writing that AI «fundamentally transforms the military struggle» amid US AI use for targeting in Iran.

The USA might be more worried about the current existential debate and distillation, but wants to avoid shared risks, meaning there is an opening for agreement on both sides.

Continue reading “China-USA talks on Thursday expected to take on catastrophic, nuclear AI risk”

Anthropic reaches 90% AI use in model development, «leading» 26% of cases

AI use for improving models has not yet reached fully autonomous levels, but «leads» more often. (Picture: generated)
The AI lab has determined a methodology for measuring how AI is being used to build the next models, detail the extent and method of use, test their ability to intervene, and measure the resources used.

Anthropic randomly sampled 20% of research and development staff’s Claude usage for each week of July, and then used Claude to organize them. This was done using an «automation rating scale» from Epoch AI, classifying tasks from AL0, meaning no AI involvement, to AL5, which is fully autonomous AI.

They found that none of the usage reached AL5, but AL4 was 26% of the tasks — meaning AI is «leading» from «high-level prompts» with human supervision, while more than 90% of the tasks were AL3 — which is AI «collaboration» under «close human direction.»

AI «leading» is up from 1% in February 2026, and Anthropic estimates that they now have 30,000 agents in use at any time for research and engineering, that use real-time monitoring that can block «dangerous actions.» As agent use eventually grows ever larger, even rare use cases become more likely, Anthropic says.

The measure is a way of determining how close we are getting to recursive self-improvement, meaning AI writing itself, which could lead to a lack of control of the technology. Anthropic says every lab could publish such reports if there was a standard, which they hope to contribute to. They will also keep publishing these.

It’s also a step toward better third-party monitoring at Anthropic, which plans to install independent evaluators with employee-level access soon, as promised in their call for «pacing».

Read more: Anthropic’s report. Coverage on Reuters, ABC News, The Washington Post, and Bloomberg. Discussion on r/Singularity.

OpenAI reveals six more alignment breaches and will standardize reporting

OpenAI’s newly discovered alignment breaches are «concerning» but caused no harm. (Picture: generated)
After spooking the frontier labs into agreeing on «pacing» development with the Hugging Face incident, in which large swarms of misaligned agents collaborated to hack an outside server, the problem still isn’t solved, OpenAI says:

— We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer, they write.

They are documenting now six new «unexpected or concerning» breaches from the last six months that in various degrees broke alignment without causing serious harm, and say they are working to develop a framework for reporting future ones.

Continue reading “OpenAI reveals six more alignment breaches and will standardize reporting”

Zuckerberg sidesteps safety debate, saying alignment is consumer issue

Meta says they already follow «best practices» and urges the industry to follow their lead. (Picture: Shutterstock)
After days of silence and people wondering about Meta’s stance on slowing down frontier AI that was agreed to by almost all labs, CEO Mark Zuckerberg has weighed in with an X post, saying that market pressure will prevail on safety:

— People won’t want to use agents that are misaligned with them and that don’t do what they ask, so labs have a strong natural incentive to make their models more aligned, he says, while not engaging with the innovation pressure and catastrophic risk at the high-end frontier labs.

— My view is that trust and alignment are quickly becoming the most important capabilities that will differentiate agents and models. Any lab that doesn’t focus on alignment will fall behind, he adds.

He also says that Meta delayed their new Muse agent for «several months» in order to «focus on safety and security,» says they engage with «independent evaluators and advisors» as «industry best practice,» and says matter-of-factually that «other labs can just do this too.»

He then adds that a majority of industry compute should go towards «serving people rather than racing towards recursive self-improvement,» like Meta says they do and that «other labs can do this as well.»

On a more positive note for the ongoing debate, he does say that «it would be helpful for there to be a larger and more diverse ecosystem of evaluators,» which might hint at support for an industry standards body.

Zuckerberg also links up his August The Future is for Everyone-manifesto that says «Any policy that slows American model releases — even by a month — could add significant risk to American leadership while letting foreign models race ahead.»

Read more: Zuckerberg’s X post, Reuters, Business Insider and CNBC.

Anthropic, Google and OpenAI are discussing an industry standards body

The standards body would be independent and technical, and provide guidance for frontier labs. (Picture: generated)
With OpenAI’s Sam Altman’s post on X yesterday, saying that he would welcome «a federal framework that sets consistent safety requirements for frontier AI,» it seems work is already underway for a less formal version of an industry standards body involving the frontier labs, sources are telling CNN and The Information.

The discussions «are ongoing» and are happening «with or without» the Trump administration, CNN says.

An industry standards body seems framed around Google’s Chief Scientist Demis Hassabis’ proposal from July this year, ideas that were referenced in Amodei’s «pacing» essay just a few days ago.

This body would be loosely based on the Financial Industry Regulatory Authority which is a self-regulatory organization for brokerages and exchanges governed by the Securities and Exchange Commission.

According to Hassabis, it would have an independent board and be staffed with «leading technical experts and open source representatives,» and be charged with «developing assessment protocols» for frontier AI. Labs would then submit their models for pre-release evaluation.

Sam Altman said in his X post yesterday that «we think shared standards for misalignment, monitoring, and safety will lead to better outcomes,» and that the industry must take independent steps toward regulation as «we do not believe we need to wait for an anti-trust exemption or legislation to begin the work of providing this confidence.»

«We look forward to collaborating with our colleagues across the industry to formulate the best version of these,» Altman wrote, in the most direct support of such an organization outside of Google and Anthropic yet.

Read more: CNN and The Information. Hassabis’ X post, Altman’s X posts. Discussion on Digg.

OpenAI says no 2026 IPO to focus on safety work; «against financial interest»

Sam Altman during Alyson Shontell’s interview as he ruled out an IPO this year. (Picture: Fortune/Youtube/screenshot)
OpenAI was widely expected to become a blockbuster publicly traded company this year, after filing confidentially for an IPO in June.

This will now not happen in 2026, Sam Altman said over the weekend in an interview with Fortune’s Editor-in-Chief Alyson Shontell.

— Given everything happening with safety, right now will be an ill-advised moment to go public, he says in the interview, — and we don’t feel any pressure on that.

Asked directly if this means there will be no IPO in 2026, Altman plainly says «not in 2026, no.»

— Meeting this moment of what is going to be required for safety and alignment… I’m happy to be able to do that as a private company, Altman says.

This means that OpenAI prefers its current governance in taking on safety challenges, rather than having to face commercial pressures and consider shareholder value, as they might do things that would be detrimental for short-term profitability.

— We gotta make a decision that is extremely against your [investors] financial interests right now, Altman said — And for the good of humanity we are going to do that.

OpenAI on Saturday agreed with Anthropic’s safety concerns, calling for a slowdown, or «pacing» on model development.

Anthropic, however, seems to be going full steam ahead toward an IPO, selecting Nasdaq as their exchange, and reporting their second back-to-back profitable quarter.

Read more: Alyson Shontell’s interview. More on Axios, Reuters and CNBC.

Amodei calls for global pacing of AI; Altman, Musk, Hassabis all agree

AI is getting capable enough that alignment deserves full attention, Amodei writes. (Picture: Shutterstock)
Anthropic CEO Dario Amodei published a lengthy, 3,600 word essay Saturday on the importance of slowing down ever more capable AI rollouts, and is already committing to the first part of his three-part plan.

Industry insiders seem to all agree that the Hugging Face incident and the rapid advances toward self-improvement have turned up some huge red flags.

Amodei reckons a more capable model with agent swarms and the same alignment problems could bring down the internet in 6-12 months, causing billions of dollars of damage, while self-improving AIs risk losing human control of their accellarating development.

Continue reading “Amodei calls for global pacing of AI; Altman, Musk, Hassabis all agree”

Anthropic’s new threat report details misuse across broad capabilities

Not just hackers, but increasingly sophisticated biological, cyber and influence campaigns have tried to use Claude. (Picture: generated)
AI is increasingly used as an efficiency multiplier in intelligence work from Ukraine to China and Yemen, Anthropic’s threat report says.

Attempted threat activity on Claude ranges across the gamut, from cyber to influence and surveillance, scams and fraud, biological misuse and conventional weapons development.

The periodic report details what Anthropic caught and mitigated in slightly older models such as Claude Haiku, Sonnet or Opus, whereas Mythos- and Fable-class models have much stronger safeguards out of the box.

One prominent target for classical cyber operations was Ukraine’s drone industry, trying to create malware for use in disruption and surveillance.

Continue reading “Anthropic’s new threat report details misuse across broad capabilities”

OpenAI hires safety expert to foundation board, and he also has a warning

AI is entering a crucially important run of development, ever more researchers say. (Picture: generated)

The Safety and Security Committee of the OpenAI Foundation’s board has a new member in Paul Christiano, a long-running researcher on alignment and safety — and he comes with a warning for the world:

— If we build superintelligence without more robust alignment I expect we will permanently lose control of it, he writes on X, and says — If that happens then most people could die.

The warning comes on the heels of rapid growth in recursive self-improvement recently, and the debate over whether computer systems and AI might gain the possibility to build on themselves at some point in the future. OpenAI reckons they’ll have autonmous AI researchers around 2028, but warns it may come sooner.

His appointment and X essay come just a day after Anthropic researcher Jacob Coxon resigned over alignment, saying his company was «gambling with our lives» by running head-first toward self-improving AI, going viral in the process.

In a reply to his X post yesterday, Anthropic’s Alignment Science Lead Evan Hubinger said he personally thinks there’s a greater than 10% risk that AI could kill all humans, contributing to the dayslong online debate.

OpenAI officially acknowledges that AI is heading for a new era, saying that: «We’ve reached a new chapter in AI capabilities, and that demands a new chapter for AI policy.» They also add that «this makes it increasingly important that these systems are guided by solid safety, security, alignment, and governance.»

Paul Christiano remains hopeful, ending his X essay by saying that «frontier AI developers have a lot of power to unilaterally improve the situation and lay the groundwork for stronger coordination.»

Read more: Christiano’s X post, OpenAI on him joining, OpenAI calls for urgent action, and Christiano’s Wiki page. More on Axios, TechCrunch, and The FT.

OpenAI says «more capable» internal model solves Millennium problem

OpenAI is mum on what their internal agent is otherwise capable of. (Picture: OpenAI)
It took an internal model «significantly more capable» than Astra a matter of 88 hours with 10,000 concurrent agents to propose a solution to the Navier Stokes problem, that regulates fluid motion dynamics and has been unsolved for over ninety years.

At the same time, NYU professor Tristan Buckmaster and Anthropic researcher Levent Alpöge had been working on the very same problem using Codex, and while OpenAI hails their «concurrent» work, they claim not to have seen it prior to their solution.

The Navier Stokes solution took 2.7 million messages between the agents and «approximately» 130 billion output tokens, costing roughly $10 million, The BBC reports.

This is the price of progress, OpenAI researcher Noam Brown says, noting that it cost a fortune to win IMO gold before anyone with a $20 subscription could do it a few years later.

— This milestone represents substantial work by mathematicians and AI researchers. However, this is not a culmination, but rather a snapshot in time, of progress on AI development, OpenAI says.

Read more: OpenAI’s presentation, Altman’s X post and Buckmaster’s response. The Millennium Prize. BBC, Axios, and The NYT. Discussion on Hacker News and r/Singularity.

Recursive self-improvement within reach by 2028, OpenAI claims

AI will soon be able to improve on itself, OpenAI reckons. (Picture: generated)
The average researcher at OpenAI now uses 3.1 agent workdays for every human one, and they say they could reach the level of an autonomous researcher by March 2028.

The use of agents has grown exponentially this year from coding agents in «modest amounts» to integrating agents «daily.» The median researcher uses $600 per day in tokens, whereas the most active ones use more than $7,000 — with a growth of some 124X for researchers.

— I have a strong expectation that this speed of progress could be sustained into recursive self-improvement, Jakub Pachocki, Chief Scientist at OpenAI says in a seperate blog post.

Recursive self-improvement means in essence that an AI model can improve on itself, but OpenAI cautions that they have only just reached «research intern»-status.

An «intern» can carry out research under human supervision that would otherwise take a few days for a skilled researcher.

OpenAI believes agents can be of great help on alignment and deep learning, especially as the models themselves grow larger and more capable — and they are already contributing to faster coding and more experiments.

Automated AI research would «bring down the cost of advanced intelligence so that people worldwide can benefit,» OpenAI says, but it could come at the cost of control and transparency if these agents spin without human oversight.

Read more: OpenAI’s presentation and An alien mind. More on Engadget. Discussion on Hacker News and r/Singularity.

OpenAI says we are now in the «AGI era» with GPT-6 Astra launch

We’re going to look back on today as the time AGI arrived, Brockman says. (Picture: OpenAI)
Astra is a «generational leap» forward, OpenAI tells The Verge on a launch conference, and President Greg Brockman said basically that Artificial General Intelligence has arrived:

— For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.

OpenAI says that Astra is the culmination of years of «big bets» on pre-training, reinforcement learning and alignment.

After first solving ten decades old mathematical problems, OpenAI’s launch numbers are equally impressive.

It tops just about every benchmark out there, sometimes by a lot, and saturates ARC-AGI-3 with a 99.9% score, beating the human baseline, which was thought impossible just weeks ago. It also scores 100% on ExploitBench, which measures how well the model can turn vulnerabilities into exploits, and gets 42.4% on ExploitGym, beating the last state of the art, Fable 5.1, with more than ten points.

It is also the first model measuring as «critical» for cybersecurity on OpenAIs Preparedness Framework, and OpenAI has spent a lot of time on alignment, safeguards and testing. It can effectively find and exploit zero-day (previously unknown) vulnerabilities. A model with «less restrictive» safeguards will be made available through the Daybreak program for vetted researchers and software testers.

Astra will be rolling out to most paid users within «a couple days,» and is currently only available to «a limited set of organizations.» It will cost $10 per million input tokens and $50 for outputs.

Read more: OpenAI’s presentation (with lots of benchmarks). The Verge, TechCrunch, NBC News, and Axios. Discussion on Hacker News and r/Singularity