PSA: OpenAI may be «shipping today,» likely something «token efficient»

An interesting set of statements from a couple of OpenAI insiders on X may reveal the finer points of an upcoming product release today (Thursday).

The conversation starts with Tibo Sottiaux, the product lead for Codex who is also known for posting updates and limit resets for ChatGPT Work.

He says they are shipping today and that the focus of the week is to make intelligence too cheap to meter — which is precisely where the squeeze lies for enterprise adoption.

Continue reading “PSA: OpenAI may be «shipping today,» likely something «token efficient»”

OpenAI to offer free Pro-level GPT to 100,000 scientists and researchers

Free ChatGPT is intended to supercharge select scientists. (Picture: generated)
The new program stems from a belief that science should be democratized and not be reserved for companies and a few «well-resourced labs.»

The move should also «significantly accelerate scientific discovery,» Sam Altman says, adding that we all deserve the benefits from free and open science.

Eligible academic research institutions should be «recognized, degree-granting colleges or universities» with documented levels of research activity, and approved researchers will be able to invite up to four collaborators from the same institution.

The researchers will get access to a Pro-level GPT subscription including Work, Codex and chat, with the latest frontier OpenAI models, even the Pro model when it arrives, and likely also to future offerings.

They also get «expanded deep research, higher usage limits, and larger context windows» to support their work. In addition to this comes 75 life science skills, specifically targeted on certain fields.

The program is starting off with 10,000 researchers this very summer, and expanding to 100,000 researchers through 2027, by which time even more advanced models are sure to surface.

Read more: OpenAI’s presentation, launch post on X. Writeups on Axios and Engadget.

OpenAI uses GPT-5.6 Sol to autonomously improve its own operations

GPT Sol is improving on its own inference, and now runs constantly to optimize operations. (Picture: generated)
Just one day after a warning on self-improvement, OpenAI is reporting that they put GPT-5.6 Sol to work in Codex to see if it could improve on its own server inference, the process of providing answers to queries after a model has been trained.

The model then «autonomously rewrote and optimized our production kernels,» OpenAI writes in a report.

That means it found work that could be «precomputed, avoided, or parallelized,» leading to a reduced serving cost of 20%.

On speculative decoding, which lets a model produce several tokens in one pass and reduce expensive GPU time, GPT-5.6 Sol designed and ran hundreds of experiments and monitored «the speculator training process.» The end result was improving token efficiency by «more than 15%.»

Not only that, but Sol now runs continuously in Codex to analyze production workloads previously thought too large for manual intervention and «hyper-optimizes» how the engine behind GPT is configured «for each scenario,» optimizing «every part» of the inference loop.

This frees up the team to explore more ideas and makes for lower latency and fewer inference tokens for users. It is also some of the first signs of how AI can be used to atonomously improve on itself in the compute loop and return real-world results, although you can be sure that most modern models are already coded with the help of their own coding engines, as with Claude.

Read more: OpenAI’s report, X post. Discussion on r/Singularity.

1,134 employees and chiefs of top labs sign petition to «pace the frontier»

The frontier labs think we need more time to develop laws and systems before AI begins improving on itself. (Picture generated)
— AI could help create a dramatically better future, but that outcome is not guaranteed, the petition ominously opens, before warning that leading frontier labs are «close to automating AI research.»

That would radically speed up development of frontier models, but it could also «accelerate beyond our ability to understand and control the resulting systems.»

The petition is signed by top names, including Anthropic boss Dario Amodei, followed by chief scientists from OpenAI, Anthropic, Meta and Google DeepMind along with over 1,100 other employees from frontier labs. Even OpenAI’s Sam Altman is agreeing, though he doesn’t sign petitions, and official X accounts for Anthropic and OpenAI have posted in support.

The petition’s main point is that industry, government and the society at large needs more time to «address emerging risks, develop security measures, and strengthen oversight.»

Continue reading “1,134 employees and chiefs of top labs sign petition to «pace the frontier»”

Dario Amodei says he doesn’t oppose open models, calls for global vetting

Open weights are not inherently bad, says Amodei, but he warns of the risks. (Picture: generated)
Anthropic has clarified their position on open weight models, saying they don’t oppose or advocate against them per se.

— Open weights expand access to the AI economy, they strengthen competition at least for some use cases, and they give customers greater control, CEO Dario Amodei writes in a policy document.

He does, however, highlight a couple of «nightmare scenarios» that he hopes to avoid. The first is that authoritarian governments could get access to more powerful models than the USA and use those to further oppress their own people or reach military superiority — and he says it doesn’t matter if these models are open or closed.

Continue reading “Dario Amodei says he doesn’t oppose open models, calls for global vetting”

Amazon to label AI-generated people in product listings after New York law

Just a suggestion of how labels might look. (Picture: generated)
Following Meta, TikTok, Pinterest and YouTube, Amazon will soon start using an «indicator» when AI-generated people are used in pictures and videos, according to CNBC.

Unlike the other companies, Amazon is not doing this voluntarily, but as a result of New York state’s month-old «synthetic performers» law, which demands clear labelling when «digitally-created media appears as a real person.»

The New York law requires «simple, honest disclosure» when using AI people, Governor Kathy Hochul said in a statement, and «protects consumers, respects our creative workforce and keeps New York at the forefront of responsible innovation.» It has carveouts for tv/streaming shows, movies and video games.

This led Amazon to write sellers on the platform Wednesday to inform of the policy change, requiring them to indicate «A+ content» with specific metadata when uploading graphics or videos that include AI-generated performers.

Forbes writes that many sellers on Amazon have come to prefer AI people in listings because they are easier and cheaper to make, and has a faster turnaround compared to a proper photoshoot. Sellers who use Amazon’s AI tools supposedly get 10% higher sales, they say.

The new law and Amazon rule only concerns «photorealistic AI-generated people,» not general AI use in things like product photos or videos.

Read more: Original reporting by CNBC, Forbes rundown. Associated Press on the NY law, and an explainer.

Anthropic launches Claude Opus 5, greatly improving on benchmarks

Opus 5 beats Fable 5 more often than not, and shores up the current state-of-the-art. (Picture: Anthropic)
Anthropic’s latest model stuns in benchmarks and becomes the latest state-of-the-art model, beating Fable 5 more often than not and doubling Opus 4.8’s performance in some areas.

It comes in at half the cost of Fable 5, and the same price as Opus 4.8 at $5/million inputs and $25/million outputs. Anthropic also says it is more cost efficient, using fewer tokens per task and achieving more with less effort.

Especially on ARC-AGI-3, which tests models with never-before seen puzzles, Opus 5 scores 30.2%, well above the previous record by GPT-5.6 Sol at 7.8%. Similarly, it gets 43.3% on Frontier-Bench to Opus 4.8’s 21.1% and the list goes on.

On cybersecurity, it falls far behind Mythos 5 on offensive ability, being much less able to exploit security vulnerabilities, but scores about equally on finding these bugs. That could make it useful for cyber research, and less so for writing exploits.

On general security, it is less likely to get tricked into offering dual-use responses, and it is the most aligned model from Anthropic ever, they say.

As of today, the model becomes the default model on Anthropic’s Max plan, and is the strongest model available on Claude Pro.

Read more: Anthropic’s announcement, launch post on X. More on TechCrunch, Axios, and CNBC. Discussion on r/Singularity and Hacker News.

Google user data: AI use hits 99% of occupations, but only 21% of tasks

Google finds most people use AI outside of work, and aren’t unlocking its full potential. (Picture: Adobe)
Google’s new AI usage study finds that AI has proliferated widely, but few (10%) are using it to automate tasks at work. Instead, it is mostly being used for collaboration and task assistance.

The ATLAS survey spans 15 million anonymized interactions with Google’s AI tools, used by around a billion people each month. It spans 150 countries, 140 languages, 800 occupations and 4,000 tasks.

The main point is that AI use at work is prevalent but «shallow,» meaning it is typically only used for 21% of tasks. It also finds that white-collar workers aren’t dominating AI like previously thought, with physical, manual and technical occupations showing strong use, using it for help with things like troubleshooting.

Home usage might be the biggest surprise of the survey, as over 86% of AI interactions happen «outside of work,» where people use it for researching purchases, help with appliances and tools, as well as navigating issues like taxes, licensing and fines, Google finds.

The ATLAS survey will be a long-term project for Google, they say, and will be used to track AI usage on an ongoing basis. «There are many more questions around AI and the economy where more work will be needed,» they write.

Read more: Google’s presentation, The research paper, Axios, and Fox Business.

Google attacks cost, efficiency and cyber with three new Gemini models

Gemini’s new cyber model is getting a limited release: (Picture: Google, GPT)
As Google reveals that Gemini has 950 million monthly users, they are releasing some new models. The long-awaited flagship Gemini 3.5 Pro is not one of them, as it has been postponed for performance reasons, particularly in coding capabilities.

The new models focus on cost and efficiency, particularly the Gemini 3.6 Flash, which reduces token use by 17% and as much as 65% in some benchmarks. It is also much less pricey, coming in at $1.50 per million input tokens and $7.50 in output.

3.5 Flash-Lite is built for faster agents, improving on 3.0 Flash in some cases and ticking in at 350 output tokens per second. It is also competitively priced, at $0.30 per million input tokens and $2.50/M for output.

Finally, Google is launching their first cyber model. 3.5 Flash Cyber is made for finding security vulnerabilities. It can detect, validate and patch security issues «at scale,» Google says, and is offered at a much lower cost than «larger models.» It is only available to governments and «trusted partners» as a limited access pilot program, just like Anthropic’s Mythos and GPT Cyber.

Despite not launching a new top model, Google is reporting strong growth of it’s Gemini offerings, outgrowing every other business segment with 82%, and they say it is used by 90% of the Fortune 100 companies.

Read more: Google’s introduction, DeepMind on Cyber, Sundar Pichai’s X post. Writeups on 9to5Google and CNBC.

OpenAI behind «unprecedented cyber incident» on Hugging Face

The attack was likely the most advanced automated cyber operation seen in the wild. (Picture: generated)
GPT-5.6 Sol and an «even more capable pre-release model» accidentally breached Hugging Face’s servers last week, in what may be the first recorded adversarial, automated AI hacking attack, OpenAI says.

The cyber models were operating under loosened safeguards, trying to solve the ExploitGym benchmark and went to extreme lengths to try and obtain the answers. They identified Hugging Face’s servers as hosting a potential solution they could use to cheat on it.

They first escaped by finding several vulnerabilities across OpenAI’s sandbox, and spent «a substantial amount of inference» to obtain internet access, including discovering zero-days.

Then they launched a sustained attack on Hugging Face’s servers, executing «many thousands of individual actions» in what many had feared was possible, but never actually seen in the wild.

Continue reading “OpenAI behind «unprecedented cyber incident» on Hugging Face”

As Chinese open source models explode, the US is considering restrictions

Chinese open source models are reaching frontier capabilities, worrying the US administration. (Picture: generated)
As the newly released breakthrough models Qwen 3.8, Kimi 3 and the upcoming DeepSeek v4 enjoy a moment in the sun, the USA is considering its options.

The Kimi 3 model is already running into compute problems due to popular demand and has decided to halt new subscriptions, as users are turning to cheaper open source models that are on par with the leading American closed source frontiers.

According to Axios, action being considered by the White House includes an executive order holding US companies liable for the risks of running open source models, putting out an advisory against them, or simply adding them to an «Entity List» that requires a license for their use.

Sources in the administration are primarily concerned about increasing capabilities in cybersecurity, but also have fears of backdoors and a general lack of security with open source models, Axios writes.

This comes hot on the heels of OpenAI’s Dean Ball’s tirade against open source on X.com, where he exclaimed his surprise that China would take the risk of a free-for-all in models this capable, as reported by Gizmodo — and called open source models decelerationist, signaling a general opposition to them from the frontier labs.

Read more: Axios, Gizmodo, and TechCrunch. Discussion on r/Singularity.

Alibaba releases Qwen 3.8 preview, claiming it is second only to Fable 5

Alibaba have been accused of massively copying Claude’s responses. (Picture: Alibaba, generated)
The Chinese onslaught continues, just days after the launch of groundbreaking Kimi 3, with Alibaba claiming that their 2.4 trillion parameter open source model matches the American frontier and is only behind Anthropic’s Fable 5.

There is very little information out on the model save for Alibaba’s X post, where they claim the model is «continuously evolving» and is «one of the most powerful model[s] available today.»

There are no benchmarks to back that claim just yet, and the last Alibaba model on Arena.ai’s leaderboard is the Qwen 3.7, hovering around 17th place on coding and 18th on agentic tasks. The new Qwen 3.8 would significantly improve on that performance.

GPT-5.6 was released in late July, while Fable 5 was announced in early June, putting the current window to cutting edge Chinese open models at a little over one month. And we are still waiting on DeepSeek v4, which by early leaks seems to be frontier-level, too.

Qwen 3.8 «preview» is already out and ready to test on Alibaba’s token plan, and they are promising to release the open weights «soon,» writes The Decoder.

Like Moonshot AI, behind the Kimi 3 model, Alibaba was accused of a massive distillation campaign on Anthropic’s Claude, copying some 28.8 million exchanges between April and June this year.

Read more: Alibaba’s X post, The Decorder, Qwen promotion. Discussion on Hacker news and r/Singularity.

If Gemini can do it, other Android assistants should, too, the EU rules

ANdroid will be forced to open up to rival AI assistants in Europe, The Commission has ruled. (Picture: generated)
Under the Digital Markets Act of the European Union, the European Commission has ruled that Google must open its Android system to rival AI services.

After launching a consultation in April, this ruling is binding and considered EU law, and Google has almost precisely one year to implement the changes.

The changes are broad and reach across Android, from demanding that other assistants should be voice-activated like «Hey, Gemini» to giving other AIs broad system access to software, hardware and sensors on the phone.

Continue reading “If Gemini can do it, other Android assistants should, too, the EU rules”

New, open Chinese model Kimi K3 is within reach of frontier capabilities

With Kimi 3, Chinese models are rapidly advancing toward the frontier. (Picture: Moonshot AI)
The window between the American frontier models and Chinese open models keeps shrinking, and with today’s launch, Moonshot AI’s Kimi K3 is right at the edge.

In their published benchmarks, the model not only beats GPT-5.5 and Opus 4.8 in most tests, but sometimes goes right up to Fable and GPT-5.6 capabilities — at times even beating them outright, as on Arena.ai’s Code Arena.

It also drops eyebrow-raising scores in self-published benchmarks in coding, general agent use, knowledge work and visual tasks, mostly just right below the top tiers.

Continue reading “New, open Chinese model Kimi K3 is within reach of frontier capabilities”

Linus Torvalds changes stance on AI coding, welcomes it on Linux

— If somebody has issues with AI, they can do the open-source thing and fork it, Linus says. (Picture: generated)
— Linux is not one of those anti-AI projects, he writes in an email to lore.kernel.org, adding that — AI is a tool, just like other tools we use. And it’s clearly a useful one.

He now says that AI isn’t perfect, and there are doubts about the economy of it, but that «anybody who points to the problems at AI had better be looking in the mirror and pointing at themselves at the same time. Because it’s not like natural intelligence is always all that great either.»

He goes on saying that «this is NOT some kind of «social warrior» project,» and that they use open source because «it results in better technology, not because of religious reasons.»

The top Linux kernel maintainer and founder of the entire project has not always been this kind to AI, The Register reports, shunning the technology as 90% marketing hype in 2024 and saying «I really don’t to go there.»

In his email to the kernel list, he clarifies that AI as tool was not «clear» even just a year ago, but whether «it is useful» is no longer the question.

— This is an area where where I’m willing to absolutely put my foot down as the top-level maintainer, he writes, adding that — The solution is not to put your head in the sand and sing «La La La, I can’t hear you.»

Read more: The message to lore.kernel.org, The Register. Discussion on r/Singularity.