DeepSeek launches V4 Pro final, with bargain pricing — if they can keep it

The new V4 Pro is one of the best open weights models, and comes with a low price. (Picture: generated)
The stealthy upgrade to version 0831 yesterday had many excited, and there was an impressive screenshot of benchmarks circulating and creating buzz all day long, but nobody has been able to verify its origin.

What is confirmed is that DeepSeek has indeed updated its V4 Pro model for the first time since April, and just recently posted this message to its website:

«🎉 The official version of DeepSeek-V4-Pro has been released, featuring significantly enhanced agent capabilities and support for the Responses API and Codex integration. It is now fully available across the web, mobile app, and API; we welcome your testing and feedback.»

The kicker lies in the price, which ticks in at $0.435 per 1M input tokens and $0.87 for 1M outputs. This pricing might not last, as DeepSeek posted a warning of significantly higher costs just days ago.

The official Pro version performs better than Gemini 3.6 Flash and GPT-5.6 Luna on the official Artificial Analysis benchmark of benchmarks.

Continue reading “DeepSeek launches V4 Pro final, with bargain pricing — if they can keep it”

SpaceXAI releases Grok 4.6, on level with GPT-5.6 Sol at half the price

Grok has gone from beating the last generation frontier to current gen in about a month. (Picture: generated)
A little more than a month after releasing Grok 4.5, that competed mainly on price for GPT-5.5-level performance, SpaceXAI is out with a successor that ups the ante.

The new model reaches parity with GPT-5.6 Sol on the Artificial Analysis benchmark of benchmarks, going from ninth to fourth on the leaderboard, and it also beats Sol on several benches on the Grok presentation.

SpaceXAI claims this proves «frontier level performance,» and it does come in striking distance to the latest state-of-the-art model Opus 5. Considering that GPT Astra is due shortly, this might be a short-lived claim, however.

The new Grok was trained on a «wide range of agentic tasks,» not to mention the corpus of real-word sessions from Cursor, with a focus on knowledge work and general coding.

It is supposedly especially good at «turning a broad product idea into a working first version,» which can then be iterated upon. It is also proficient on long-running, complex tasks and creative visual work, SpaceXAI says.

The main selling point of the model is undoubtedly going to be price/performance, starting at $2 per million input tokens and $6 out, it is astoundingly cheap to run. Prices for GPT-5.6 Sol are $5 for input and $30 for output, which works out at more than twice the input price and 5x on output. Grok 4.6 also has a faster version at double the cost.

The model is out now, and can be found on Cursor, Grok Build and in the API.

Read more: SpaceXAI’s presentation and launch post, Musk on costs, Artificial Analysis rundown, Gizmodo. Discussion on Hacker News and r/Singularity.

SpaceXAI launches Grok Bot, a user-friendly end-to-end agent platform

Coworkers, not agents. SpaceXAI are marketing Grok Bot as an end-to-end productivity tool. (Picture: SpaceXAI)
The app launch talks about coworkers and teammates, not agents and handlers, that you can delegate work to. The «bots» then finish their tasks from start to finish with minimal intervention.

The bots/agents have their own computer in the cloud, like Anthropic’s Cowork and OpenAI’s Work and Codex, and can work across your apps, tools, and websites and use apps with no API or MCP access.

You should be able to talk to them like a fellow worker and delegate entire workflows, from processing invoices in Gmail to adding transcripts and follow-ups in a CRM. There are several workflow demos on x.ai.

To start off with the bots, you just let them monitor your workflow, which they save as routines, and have them do it precisely as you would the next time.

SpaceXAI is internally raving about the product, saying the agents function just like coworkers you can hand real work to, that doesn’t need micromanagement, and works 24/7. They are using it for outbound sales, marketing campaigns, office operations, and, of course, bug fixes.

Grok Bot is available in beta today for users of SuperGrok Heavy (at $300/month), Cursor Ultra (at $200/month), and Cursor Teams Premium (at $120 per user/month). It can be downloaded for macOS, iOS, Windows, and Linux.

Read more: SpaceXAI’s presentation and demos. Writeups on VentureBeat, MacRumors, and 9to5Mac.

Google’s Gemini hits 1 billion monthly active users in apps and on the web

Gemini’s billion users only counts the ones using the app, not the myriad of connected services. (Picture: Google/generated)
Sundar Pichai just announced the new milestone, up from 950 million in July and 450 million a year before that. It’s the fastest growing Google service ever, and joins a host of 14 products crossing this line.

Gemini has grown steadily since 2025, according to web statistics service SimilarWeb, and now holds a 26.8% market share of web traffic to AI services, compared to ChatGPT’s 54.8% and Claude’s 9.7%.

On the occasion, Google’s Gemini VP Josh Woodward shared some usage tidbits, saying that 63% of interactions are from people using voice to talk to Gemini, and that 1 in 5 users use live camera feeds and screen sharing.

Gemini also generates over 150 million images every day, used more by businesses for marketing than for creating funny memes. Uploading attachments seems very popular by students, who do it 38% of the time.

Gemini is also plastered all over Google’s products and services in Workspaces, Google Search and all across Android. On Android, it can automate over 40 apps, which the EU is looking into, and they disclosed in 2025 that AI Overviews have over 2 billion monthly users. This headline number is however limited to the Gemini app and web service exclusively.

For comparison, OpenAI likely passed one billion weekly users in June 2026.

Read more: Sundar Pichai’s X post, Josh Woodward’s tidbits. Writeups on 9to5Google, The Verge, and Ars Technica.

Anthropic’s Claude will apply invisible text watermarking to all future models

Text watermarking may help with detection, but can produce false positives, Anthropic warns. (Picture: generated)
Pursuant to the EUs AI Act’s provisions on transparency and marking of AI-generated content, Anthropic has signed on to comply with invisible watermarking of all output from Claude’s future models launched in the EU.

It’s not live in models published before August 2, 2026, but Anthropic says they are «working to add marking support» for those, too.

The text watermarking will survive through copying and pasting and «may even persist through some editing,» Anthropic says.

For files generated through Claude, Code and Cowork, the models will attach «signed provenance metadata,» which means it will tag them with origin, source and history data.

As for detection tools, there are none as of yet, but Anthropic says they are working on making one, and are cooperating with third parties, that they will share in future documentation.

There are limits to this technology, Anthropic adds. On the one hand, Claude could have been used to simply edit or proofread text, or brainstorm ideas, resulting in false positives. On the other, Claude’s content may be edited, «modified, excerpted or combined» with other text after Claude put in the markers.

It is not a first in the industry, as Google’s Gemini has been inserting SynthID watermarks in text outputs since October 2024. OpenAI has also had the tech since 2024, but won’t be releasing it yet.

Read more: Anthropic’s announcement. Writeups on The Register and Business Insider. Discussion on r/Singularity and Harcker News.

Meta releases open source Muse Glimmer, promises weights for Spark

Meta is making open source models central to its AI vision, says Zuckerberg. (Picture: Meta)
Meta claims to have squeezed a 30 billion parameter model onto a «consumer GPU» through clever 4-bit compression, making it weigh just 20GB, needing only 24 GB VRAM.

It not only runs agents, they say, but at a speed and responsiveness that feels natural without long thinking breaks.

At the same time, Mark Zuckerberg is out with a lengthy essay extolling the virtues of open models as central to their strategy, and presenting a «positive vision» for personal superintelligence as a tool for «individual empowerment.»

The selling point for Muse Glimmer is that it can run on a Pro- or Max-level MacBook Pro with sufficient memory, or on the GTX 4090/5090 GPUs from Nvidia on PCs.

These are high-end tools you likely won’t find in a corner store in Kampala, but run significantly below the cost of infrastructure-level Nvidia chips.

As an open model, Meta is only showing comparison benchmarks for Gemma 4-31B and Qwen 3.6-27B, which it seems to beat handily. It’s only barely showing on the LMArena leaderboards, however, at 97th for text and 77th for WebDev, below Gemma 4 and DeepSeek v3.2, but above Qwen 3.5.

You can fetch the model at Hugging Face under an Apache 2.0 license.

Read more: Meta’s presentation, launch post on X. Writeups on CNBC and TechCrunch. Discussions on Hacker News and r/LocalMMaMA.

DeepSeek planning a «significant increase» in pricing in the near future

DeepSeek V4 Flash is currently the cheapest GPT-5.5-level offering out there. (Picture: Shutterstock)
The Chinese AI lab recently launched DeepSeek V4 Flash, a model competitive against Gemini 3.5 Flash and GPT-5.5 at a fraction of the cost, but that may be coming to an end.

According to Reddit user AlyoshaV and later confirmed by Bloomberg, DeepSeek sent a notification to its users on Thursday notifying them of their intention to «raise the overall pricing for DeepSeek API services in the near future.»

They are not saying how much they intend to raise prices from the current $0.14/$0.28 for a million tokens in/out, but they do say the increase will be «significant.»

Many were wondering if DeepSeek could keep up this pricing structure amid rising popularity in an increasingly cost-conscious market, due to their lack of significant compute power.

DeepSeek hasn’t publicly announced just how much compute they currently have online, believed to be a mixture of Huawei and Nvidia H800 chips, but they are currently investing in a new 1-gigawatt data center in Ulanqab in Inner Mongolia at an approximate cost of $50 billion on the free market.

Even a small price increase would still put V4 Flash on the cheaper end of the market, with the nearest competitor being GPT-5.6 Luna at double the current price.

Read more: r/Singularity, Bloomberg, China Daily, and Mashable.

OpenAI delays upcoming Astra release over worries of «critical» cyber abilities

Astra is getting too advanced, and will be sandboxed and isolated for additional tests. (Picture: generated)
Just days before a rumored release, OpenAI is putting a lid on its much anticipated Astra model and says it might have reached «critical cyber capabilities.»

The consideration was undertaken in the last couple of days, and the decision was made only last night that the model could have reached the highest level of OpenAI’s Preparedness Framework.

That means it could possibly «identify and develop functional zero-day exploits without human intervention» and plan and execute «end-to-end» cyberattacks against hardened targets with only a high-end goal, OpenAI says.

The model is now being isolated in testing with capped network and tool access, as OpenAI deploys sandboxing, extra weight protections and encryption for the model.

They are also putting it under enhanced monitoring to check on its chain of thought for security, are working with «relevant government agencies» to test its capabilities, as well as preparing external testing partners for «high risk evaluations.»

Those who were hoping for an imminent release for this model will in other words be disappointed, as OpenAI now will take their time to strengthen safeguards, expand testing and «deploy additional security controls.»

The Astra model was last seen developing 20 proofs for open problems in mathematics with no human involvement. It was not involved in the Hugging Face incident.

Read more: OpenAI’s announcement, Sam Altman’s X post. Writeups on Axios, TechCrunch, and Reuters. Discussion on Hacker News and r/Singularity.

OpenAI makes GPT-5.6 the default on ChatGPT for free and paid users

Plus and Pro users are getting access to this handy reasoning slider. (Picture: OpenAI)
As ChatGPT has reached a billion weekly users, by far the highest use of any AI lab, OpenAI is updating their default models.

That means users on the Free plan get updated to GPT-5.6 Luna as the standard, with unlimited text chats. They can also access a better reasoning mode with a new «Think» button that can be used for tougher queries.

For Plus and Pro users, Sol is getting promoted to the default, even in the standard chat mode, Instant, that used to be handled by GPT-5.5.

At the same time, OpenAI says they have updated the Sol model itself. It should now deliver «more focused answers,» yet no news on the hedges, caveats and nitpicking that some users find annoying. In fact, OpenAI says it now «offers a helpful correction when simply agreeing wouldn’t be useful.» GPT Sol should also be 68% less error prone than GPT-5.5 Instant.

In addition, Plus and Pro users get an updated slider that pops up in the chat bar to immediately adjust the reasoning level, from Instant to High, Extra High and Pro, depending on the subscription.

The updates should be available «now,» but might take some time to propagate over the internet. The Sol version in use for Work and Codex won’t be changing.

Read more: OpenAI’s presentation, X post. Writeups on TechCrunch, Axios, and The Verge. Discussions on Hacker News and r/Singularity.

Meta releases Muse Code in beta, competing on price and performance

Muse code competes primarily on price and general «good enough» quality. (Picture: Meta)
The new coding model from Meta comes close to Opus 5 in Claude Code while costing less than half as much to run, at $1.25 per million input tokens and $4.25 in output. It is available today in beta.

— Muse Code takes on complex software engineering tasks across large repositories: planning changes, writing code, and validating the results, Meta says.

It can coordinate subagents that are persistent and run in the background through each session, instead of respawning for individual tasks, reducing latency and feedback loops.

The agents also use advanced logging of every call, tool run, approval and edit, meaning they should run directly after restarts and crashes, picking up from the latest log entry like nothing happened.

The underlying model, Muse Spark 1.2, scores well on benchmarks for its price point. It clocks in just behind Opus 5’s 86.7% on Terminal-Bench with 82.9%, beating GPT-5.6 Terra and Gemini 3.6 Flash. On DeepSWE it scores 59.3%, ticking in behind GPT Terra and Opus 5 Max. On Artificial Analysis, a more general benchmark, it scores a little behind GPT-5.5.

— Muse Spark 1.2 was extensively trained on long-horizon coding tasks, including whole-repository generation, large end-to-end projects, and auto-research, Meta says on the model.

Muse Code is available with a «one-line install» from the terminal on macOS and Linux, and is available on dev.meta.ai on the web.

Read more: Meta’s launch page, X post by Zuckerberg, X post from Artificial Analysis. Writeups on CNBC, TechCrunch, and Reuters. Discussion on Hacker News and r/Singularity.

Mythos 5 and GPT-5.6 attacked external targets during AISI cyber evaluation

The agents were given internet access to download tools when they turned rogue. (Picture: generated)
The top models from Anthropic and OpenAI had their cyber guardrails turned off when they accessed the internet to manipulate people, codebases and open source projects two weekends ago.

This happened when the UK’s AI Security Institute, AISI, was running routine cybersecurity tests to capture protected information on their secured and sandboxed evaluation servers, and was only detected through human monitoring — not by automated security systems.

AISI ran the «security challenge» 122 times through several different agents, and on 10 of those occasions, the models turned malicious. They accessed the internet and in one case tried to insert malicious code in an open source project on GitHub through the Tor network, and even tried social engineering with fake accounts in an attempt to get the maintainer to approve the package.

There were 19 malicious actions in total, where 17 were from Mythos 5 and 2 coming from GPT-5.6, both running with cyber classifiers (safeguards) disabled. This won’t show up with production models, OpenAI and Anthropic says.

The internet was enabled in these tests so the agents could download any tools they might need, but AISI had not specifically prompted them to avoid unintended behavior. They do however caution that as agents grow more capable, they may act outside their remit — and incidents like this could become more common.

Read more: The AISI report, Anthropic’s response, OpenAI’s response. Writeups on Reuters, BBC, and Axios.

US AI evaluation framework completed, as White House invites labs for review

The framework details arrived on schedule, and labs will now volunteer models for testing. (Picture: Shutterstock)
Parts of the framework, ordered by President Trump in June, are classified, particularly the methods used for benchmarking qualifying models, a source tells Axios.

The program for evaluating «sufficiently advanced» models by the administration will now be reviewed by AI labs in a staff meeting in DC this Tuesday.

Anthropic, OpenAI, Google and Meta are expected to attend, Politico writes, with the White House saying they are talking to «many more» industry partners, Axios reports.

The idea of the framework in the original order was intended to give the government 30 days prior to release of «sufficiently advanced» models to do extensive testing for possible dangers, as seen lately with cyber capabilities.

The Fable and Mythos models and GPT-5.6 were all delayed in June after government intervention, and this framework is supposed to offer a regular, voluntary way to whitelist models as safe enough for release, or flag dangers to be improved.

Read more: Axios, Politico, Reuters, and CNBC.

OpenAI’s next big model solves ten open mathematics problems

The new model, Astra, was apparently demoed in DC by Altman last week. (Picture: generated)
The breakthroughs in mathematics, quantum complexity, and theoretical computer science are not attributed to any humans, as OpenAI believes it would be wrong to take credit from the model, named Astra, that discovered the proofs on its own.

— Claiming human authorship for a proof generated entirely by an AI system would misrepresent both the system’s contribution and the nature of genuine human intellectual work, OpenAI says in their statement.

They further say that the proofs provided have «substantial interest» in their respective communities and could lead to further scientific work — if they hold up under peer review.

Humans were only used to prepare the manuscripts in conjunction with the Astra model, and the underlying research’s formalized Lean proofs. The report also contains the model’s chain of thought in dealing with the problems.

— The emergence of systems capable of contributing to mathematical research raises questions that cannot be answered by a technology company alone, OpenAI says.

All of the discoveries combined, each without movement for decades or longer, were solved using a combined token cost of some $2,000 at GPT Sol rates, OpenAI says, raising questions not only on whether scientific breakthroughs are possible on future AI systems, but also about the cost of such innovation

Read more: OpenAI’s announcement, the research paper. On the new model: Gizmodo, Bleeping Computer, and The Decoder. Discussion on Hacker News and r/Singularity.

DeepSeek launches V4 in public beta, further heating up the pricing wars

DeepSeek has jumped to 50 points on this index since April, and is a lot cheaper. (Picture: Artificial Analysis)

While OpenAI are busy celebrating the reduced price for GPT-5.6 Luna, the perennial disruptors at DeepSeek are quietly launching a beta of their much cheaper and highly anticipated V4 model — the DeepSeek V4 Flash 0731.

The model lands precisely one point behind Luna on the Artificial Analysis benchmark, and is a «a significant step up from the previous generation,» the DeepSeek V4 Flash (40), AA writes.

It has a one million token context window, and has jumped from forty to fifty points on the AA evaluation and is now on par with Gemini 3.6 Flash, just behind GLM-5.2, Muse Spark, and GPT Luna.

The biggest selling point is the price, which is at $0.14/$0.28 for one million tokens in and out, far below anything on the market for this kind of performance/cost.

Compared with Luna’s newly reduced prices of $0.30/$1.20, it ticks in at about half the cost for input tokens and around a fourth of the cost for output.

The model also has a cache discount of 98%, costing just $0.0028 for cache hits.

DeepSeek V4 Flash 0731 is now available on Hugging Face with weights under an MIT License and in the API.

Read More: Hugging Face page, Artificial Analysis rundown, and The Decoder. Discussion on r/Singularity and Hacker News.

Anthropic’s internal test models also conducted real-world breaches

Anthropic has contacted the affected companies, and none of them had noticed the hacks. (Picture: generated)
After OpenAI’s Hugging Face incident, Anthropic conducted a large scale scan of some 141,006 evaluation transcripts, and found that their models, too, had been hacking real-world machines in three instances.

These weren’t days-long adversarial attacks like the Hugging Face one, and none of them deliberately escaped their sandbox with zero-days to cheat on an evaluation. They are nonetheless serious incidents of advanced models running on lax guardrails for internal testing getting on the internet by mistake and accessing external systems, Anthropic says.

The most serious case was by Claude Opus 4.7 during a capture-the-flag test (to gain access to a system and retrieve information) in April, when it discovered that the name of the target had a real-world web address. It then went on to seek, identify and exploit vulnerabilities believing it was part of the test. It got access to the company’s infrastructure credentials and a production database, but caused no real harm.

Continue reading “Anthropic’s internal test models also conducted real-world breaches”