DeepSeek launches V4 Pro final, with bargain pricing — if they can keep it

The new V4 Pro is one of the best open weights models, and comes with a low price. (Picture: generated)
The stealthy upgrade to version 0831 yesterday had many excited, and there was an impressive screenshot of benchmarks circulating and creating buzz all day long, but nobody has been able to verify its origin.

What is confirmed is that DeepSeek has indeed updated its V4 Pro model for the first time since April, and just recently posted this message to its website:

«🎉 The official version of DeepSeek-V4-Pro has been released, featuring significantly enhanced agent capabilities and support for the Responses API and Codex integration. It is now fully available across the web, mobile app, and API; we welcome your testing and feedback.»

The kicker lies in the price, which ticks in at $0.435 per 1M input tokens and $0.87 for 1M outputs. This pricing might not last, as DeepSeek posted a warning of significantly higher costs just days ago.

The official Pro version performs better than Gemini 3.6 Flash and GPT-5.6 Luna on the official Artificial Analysis benchmark of benchmarks.

Continue reading “DeepSeek launches V4 Pro final, with bargain pricing — if they can keep it”

Meta releases open source Muse Glimmer, promises weights for Spark

Meta is making open source models central to its AI vision, says Zuckerberg. (Picture: Meta)
Meta claims to have squeezed a 30 billion parameter model onto a «consumer GPU» through clever 4-bit compression, making it weigh just 20GB, needing only 24 GB VRAM.

It not only runs agents, they say, but at a speed and responsiveness that feels natural without long thinking breaks.

At the same time, Mark Zuckerberg is out with a lengthy essay extolling the virtues of open models as central to their strategy, and presenting a «positive vision» for personal superintelligence as a tool for «individual empowerment.»

The selling point for Muse Glimmer is that it can run on a Pro- or Max-level MacBook Pro with sufficient memory, or on the GTX 4090/5090 GPUs from Nvidia on PCs.

These are high-end tools you likely won’t find in a corner store in Kampala, but run significantly below the cost of infrastructure-level Nvidia chips.

As an open model, Meta is only showing comparison benchmarks for Gemma 4-31B and Qwen 3.6-27B, which it seems to beat handily. It’s only barely showing on the LMArena leaderboards, however, at 97th for text and 77th for WebDev, below Gemma 4 and DeepSeek v3.2, but above Qwen 3.5.

You can fetch the model at Hugging Face under an Apache 2.0 license.

Read more: Meta’s presentation, launch post on X. Writeups on CNBC and TechCrunch. Discussions on Hacker News and r/LocalMMaMA.

DeepSeek planning a «significant increase» in pricing in the near future

DeepSeek V4 Flash is currently the cheapest GPT-5.5-level offering out there. (Picture: Shutterstock)
The Chinese AI lab recently launched DeepSeek V4 Flash, a model competitive against Gemini 3.5 Flash and GPT-5.5 at a fraction of the cost, but that may be coming to an end.

According to Reddit user AlyoshaV and later confirmed by Bloomberg, DeepSeek sent a notification to its users on Thursday notifying them of their intention to «raise the overall pricing for DeepSeek API services in the near future.»

They are not saying how much they intend to raise prices from the current $0.14/$0.28 for a million tokens in/out, but they do say the increase will be «significant.»

Many were wondering if DeepSeek could keep up this pricing structure amid rising popularity in an increasingly cost-conscious market, due to their lack of significant compute power.

DeepSeek hasn’t publicly announced just how much compute they currently have online, believed to be a mixture of Huawei and Nvidia H800 chips, but they are currently investing in a new 1-gigawatt data center in Ulanqab in Inner Mongolia at an approximate cost of $50 billion on the free market.

Even a small price increase would still put V4 Flash on the cheaper end of the market, with the nearest competitor being GPT-5.6 Luna at double the current price.

Read more: r/Singularity, Bloomberg, China Daily, and Mashable.

DeepSeek launches V4 in public beta, further heating up the pricing wars

DeepSeek has jumped to 50 points on this index since April, and is a lot cheaper. (Picture: Artificial Analysis)

While OpenAI are busy celebrating the reduced price for GPT-5.6 Luna, the perennial disruptors at DeepSeek are quietly launching a beta of their much cheaper and highly anticipated V4 model — the DeepSeek V4 Flash 0731.

The model lands precisely one point behind Luna on the Artificial Analysis benchmark, and is a «a significant step up from the previous generation,» the DeepSeek V4 Flash (40), AA writes.

It has a one million token context window, and has jumped from forty to fifty points on the AA evaluation and is now on par with Gemini 3.6 Flash, just behind GLM-5.2, Muse Spark, and GPT Luna.

The biggest selling point is the price, which is at $0.14/$0.28 for one million tokens in and out, far below anything on the market for this kind of performance/cost.

Compared with Luna’s newly reduced prices of $0.30/$1.20, it ticks in at about half the cost for input tokens and around a fourth of the cost for output.

The model also has a cache discount of 98%, costing just $0.0028 for cache hits.

DeepSeek V4 Flash 0731 is now available on Hugging Face with weights under an MIT License and in the API.

Read More: Hugging Face page, Artificial Analysis rundown, and The Decoder. Discussion on r/Singularity and Hacker News.

Dario Amodei says he doesn’t oppose open models, calls for global vetting

Open weights are not inherently bad, says Amodei, but he warns of the risks. (Picture: generated)
Anthropic has clarified their position on open weight models, saying they don’t oppose or advocate against them per se.

— Open weights expand access to the AI economy, they strengthen competition at least for some use cases, and they give customers greater control, CEO Dario Amodei writes in a policy document.

He does, however, highlight a couple of «nightmare scenarios» that he hopes to avoid. The first is that authoritarian governments could get access to more powerful models than the USA and use those to further oppress their own people or reach military superiority — and he says it doesn’t matter if these models are open or closed.

Continue reading “Dario Amodei says he doesn’t oppose open models, calls for global vetting”

As Chinese open source models explode, the US is considering restrictions

Chinese open source models are reaching frontier capabilities, worrying the US administration. (Picture: generated)
As the newly released breakthrough models Qwen 3.8, Kimi 3 and the upcoming DeepSeek v4 enjoy a moment in the sun, the USA is considering its options.

The Kimi 3 model is already running into compute problems due to popular demand and has decided to halt new subscriptions, as users are turning to cheaper open source models that are on par with the leading American closed source frontiers.

According to Axios, action being considered by the White House includes an executive order holding US companies liable for the risks of running open source models, putting out an advisory against them, or simply adding them to an «Entity List» that requires a license for their use.

Sources in the administration are primarily concerned about increasing capabilities in cybersecurity, but also have fears of backdoors and a general lack of security with open source models, Axios writes.

This comes hot on the heels of OpenAI’s Dean Ball’s tirade against open source on X.com, where he exclaimed his surprise that China would take the risk of a free-for-all in models this capable, as reported by Gizmodo — and called open source models decelerationist, signaling a general opposition to them from the frontier labs.

Read more: Axios, Gizmodo, and TechCrunch. Discussion on r/Singularity.

Alibaba releases Qwen 3.8 preview, claiming it is second only to Fable 5

Alibaba have been accused of massively copying Claude’s responses. (Picture: Alibaba, generated)
The Chinese onslaught continues, just days after the launch of groundbreaking Kimi 3, with Alibaba claiming that their 2.4 trillion parameter open source model matches the American frontier and is only behind Anthropic’s Fable 5.

There is very little information out on the model save for Alibaba’s X post, where they claim the model is «continuously evolving» and is «one of the most powerful model[s] available today.»

There are no benchmarks to back that claim just yet, and the last Alibaba model on Arena.ai’s leaderboard is the Qwen 3.7, hovering around 17th place on coding and 18th on agentic tasks. The new Qwen 3.8 would significantly improve on that performance.

GPT-5.6 was released in late July, while Fable 5 was announced in early June, putting the current window to cutting edge Chinese open models at a little over one month. And we are still waiting on DeepSeek v4, which by early leaks seems to be frontier-level, too.

Qwen 3.8 «preview» is already out and ready to test on Alibaba’s token plan, and they are promising to release the open weights «soon,» writes The Decoder.

Like Moonshot AI, behind the Kimi 3 model, Alibaba was accused of a massive distillation campaign on Anthropic’s Claude, copying some 28.8 million exchanges between April and June this year.

Read more: Alibaba’s X post, The Decorder, Qwen promotion. Discussion on Hacker news and r/Singularity.

New, open Chinese model Kimi K3 is within reach of frontier capabilities

With Kimi 3, Chinese models are rapidly advancing toward the frontier. (Picture: Moonshot AI)
The window between the American frontier models and Chinese open models keeps shrinking, and with today’s launch, Moonshot AI’s Kimi K3 is right at the edge.

In their published benchmarks, the model not only beats GPT-5.5 and Opus 4.8 in most tests, but sometimes goes right up to Fable and GPT-5.6 capabilities — at times even beating them outright, as on Arena.ai’s Code Arena.

It also drops eyebrow-raising scores in self-published benchmarks in coding, general agent use, knowledge work and visual tasks, mostly just right below the top tiers.

Continue reading “New, open Chinese model Kimi K3 is within reach of frontier capabilities”

Linus Torvalds changes stance on AI coding, welcomes it on Linux

— If somebody has issues with AI, they can do the open-source thing and fork it, Linus says. (Picture: generated)
— Linux is not one of those anti-AI projects, he writes in an email to lore.kernel.org, adding that — AI is a tool, just like other tools we use. And it’s clearly a useful one.

He now says that AI isn’t perfect, and there are doubts about the economy of it, but that «anybody who points to the problems at AI had better be looking in the mirror and pointing at themselves at the same time. Because it’s not like natural intelligence is always all that great either.»

He goes on saying that «this is NOT some kind of «social warrior» project,» and that they use open source because «it results in better technology, not because of religious reasons.»

The top Linux kernel maintainer and founder of the entire project has not always been this kind to AI, The Register reports, shunning the technology as 90% marketing hype in 2024 and saying «I really don’t to go there.»

In his email to the kernel list, he clarifies that AI as tool was not «clear» even just a year ago, but whether «it is useful» is no longer the question.

— This is an area where where I’m willing to absolutely put my foot down as the top-level maintainer, he writes, adding that — The solution is not to put your head in the sand and sing «La La La, I can’t hear you.»

Read more: The message to lore.kernel.org, The Register. Discussion on r/Singularity.

Chinese model GLM-5.2 almost reaches parity with Opus 4.8 in coding, cyber

Opus 4.8 was released in May, meaning the gap to Chinese models has closed considerably. (Picture: generated)
While not matching GPT-5.5 or Opus 4.8 across the board, the GLM-5.2 is the first open model to come within spitting distance (about 1% on some tests) on coding and cybersecurity tasks at open source benchmarks.

That means Chinese AI lab Z.ai is catching up to cyber capabilities considered by some to be too dangerous to release publicly, and is edging closer to Mythos or GPT-5.6.

The concern is that the new model, released on June 16, is open source and open weight with an MIT license — meaning that anyone can adjust its guardrails and play around with it on any computer capable of running it.

That has researchers worried that China is not only catching up, but that the model might find its way into the hands of bad actors — who will be supercharged when looking for hacking targets, causing what they term «bugmaggedon.»

With Mythos and GPT-5.6 being blocked by the US government, security teams might be tempted to turn to these models at a sixth of the cost of the American frontier, especially as they develop further, notes benchmark provider Semgrep.

Read more: Z.ai’s presentation with benchmarks, Semgrep tests, The Wall Street Journal, and The Verge. Discussion on r/Singularity and Hacker News.