The proposed new interface standard will not only make it easier to coordinate and control things like factory floors and lab facilities, but it will also infuse AI agents into the process, allowing for scripting — and adding natural language operations and reasoning.
This could potentially be a huge boon for scientific research and real-world labs, where Claude might one day be able to run end-to-end experiments and manipulate instruments, greatly accelerating the process.
Hugging Face had been looking to raise money and was considering a sale lately. (Picture: generated)Nvidia’s revenues were up 106% to $96.22 billion last quarter, while predicting a 70% increase in the next fiscal year, so agreeing to acquire Hugging Face for $12.9 billion might seem like pocket change.
It is, however, one of Nvidia’s largest acquisitions, Reuters notes, after the buyout was first reported by The Information.
Hugging Face is the «face» of the open source AI movement and maintains a repository of almost all available models, making this a significant infrastructure investment. AI labs treat publishing their weights on it as their official release.
Nvidia is of course not new to Hugging Face or open source, having as good as bought the OSS AI lab Poolside earlier this week to build their own models. They also invested in a $235 million funding round for Hugging Face in 2023 that valued it at $4.5 billion, and tried to invest $500 million in 2025 at a $7 billion valuation, according to The Financial Times.
It appears that Nvidia is somewhat hedging their bets and is increasingly investing in open source, as the frontier AI labs are increasingly developing their own chips, Reuters writes. Nvidia remains a significant investor in closed source providers, though.
The new chip will be deployed at scale within the year. (Picture: OpenAI)Using the measure of performance per watt rather than throughput per second, OpenAI ran three open source models, GPT-OSS, DeepSeek R1 and Kimi K2.5 1T, through their new processor on the InferenceX benchmark from SemiAnalysis.
They found that the processor is wicked fast compared with «leading chips,» which means Nvidia’s offerings, specifically the Grace Blackwell 200 which they show getting trounced in some tests.
The benchmark found that Jalapeño delivers 1.9x more tokens per second when measured per watt on peak efficiency, and is 17.8x faster on higher token density, which translates to raw performance in handling requests.
They also found that end-to-end latency was between 1.7x and 3.4x lower depending on the model, meaning the user will spend less time waiting for responses.
These two measures combined are key, as most chips have to make tradeoffs between latency and throughput, and few can be good at both, OpenAI hardware vice president Richard Ho tells The Verge.
Jalapeño was developed in record time — 9 to 16 months — assisted by AI, and OpenAI says their upcoming Astra model is already busy working on the second generation, which they say is «in deep development.»
OpenAI will deploy and «operate Jalapeño at scale» within their compute infrastructure «by the end of the year.»
Nothing chews through inference tasks as fast as Nvidia’s Groq. By far. (Picture: Nvidia/generated)The never-before-seen feat was achieved running Google’s Gemma 4 31B through Artificial Analysis’ standard tests, and is way ahead of anything on the market.
The only comparable score is that of OpenAI’s GPT-Sol running on Cerebras chips, which achieved 750 tokens per second earlier in August. They called this «Ultrafast mode.»
Nvidia further says that it achieved this output score while maintaining hundreds of thousands of context tokens, and that Groq is some 34X faster than today’s quickest hardware in time to generate 5,000 tokens.
A Groq chip pairs 500 MB of high speed SRAM on die, directly next to the chip, that delivers 150 TB/s throughput. They come in racks of 256 chips stacked together for a total of 40 petabytes of memory bandwidth.
While they are great for inference tasks, GPUs will still be the workhorse of AI data centers, as they can handle both training and later inference — but Jensen Huang of Nvidia recommends setting aside 25% of data center space for the new Groq chip racks, according to CNBC.
The new test scores come as Nvidia is announcing that Groq 3 chips are now in production with Samsung and are generally available.
Even a $4 trillion company is not immune from RAMageddon. (Picture: generated)Nvidia has told its largest customers to expect a price increase for its AI systems of more than 15%, Bloomberg reports.
The hikes come amid soaring prices on memory, components and storage as AI buildouts create unprecedented demand in the market.
The systems involved will be based on both Vera Rubin and Grace Blackwell, and prices will depend on both chip and memory configurations. Increases are expected early next year, writes Reuters.
At the same time, Nvidia is pushing harder into making its own AI services, announcing what is «not an acquisition» and «not an acquihire» — before doing both to AI startup Poolside.
Nvidia will be paying $6 billion to the company to license its software, and is concurrently offering jobs to 109 of its staff. On top of that comes a straight-up investment of $1 billion at a $12 billion valuation.
ChatGPT can now send messages on your behalf, but be careful with the permissions. (Picture: generated)ChatGPT with Messages works in Work and Codex but not in regular chats. It can work across apps on your computer, so you can ask it to check the calendar for which days you are free for dinner and send it in a message to anyone in your Contacts.
The plugin can also search and analyze your messages and give you a list of, say, who you exchange the most messages with, which spam messages you can delete, or what messages need follow-ups.
OpenAI does say to be careful with permissions, as the plugin is required by default to get your approval before it sends any messages in your name. Without this setting, things can quickly get out of hand, so they recommend you keep it on.
— Persistent approval removes your final chance to review a message before ChatGPT sends it as you. Use it only when you accept that risk, OpenAI warns in its instructions.
The plugin runs locally on the user’s Mac and «doesn’t create an index of someone’s messages,» according to TechCrunch.
ChatGPT with Messages is free and available on all plans in the ChatGPT desktop app for macOS on Apple Silicon, and it can read and reply to anything Messages can receive. It does not, however, work the other way — and won’t let you interact with ChatGPT through SMS.
OpenRouter says they are vendor neutral and builds on the same principles as Stripe. (Picture: OpenRouter, generated)OpenRouter will continue as normal after the deal, supporting token- and task-optimized routing between models and keeping their neutrality.
Stripe is a huge fintech darling that routes some $1.9 trillion in payments each year and has a revenue of $5.1 billion from clients like Ford, Spotify, OpenAI and Anthropic. The still privately held company has attracted early investments from Elon Musk and Peter Thiel, according to Wikipedia, and they are also trying to buy PayPal, Axios reports.
OpenRouter had raised $164 million in May 2026 at a valuation of $1.3 billion, so the Axios report of an $8 billion price is at a premium as well as highlighting their stellar growth since their establishment in 2023, just as the AI market started emerging.
— Tokens are the central currency for companies building with AI, and it’s clear that the real-world economic potential will depend on making good use of scarce compute resources, says Patrick Collison, co-founder and CEO of Stripe in their press release, hinting at how price-sensitive enterprise AI usage has become.
Earlier this year, Stripe were giddish on AI, declaring that the start of 2026 also marked «the beginning of the singularity,» according to an investor letter published by Eric Newcomer, as they reached customers in 88% of the Forbes AI 50 list, according to Axios.
— As models become more capable, the risks associated with developing and testing them internally also grow. Our standards for monitoring, alignment, and security must stay ahead of those risks, OpenAI writes.
Combined with rapid progress in their internal research, OpenAI says they have implemented a two week pause in training future models, while the «largest planned frontier reinforcement learning run» is put on hold.
They will now implement stronger sandboxing for certain code execution, better network isolation and internet caps, and pursue continuous testing for security, alignment, deception and reward hacking, with a thirty minute warning system for any adverse incident.
The AI lab expects most of this safety training to be handled by their own models in the future, greatly expanding their scope while also reducing the time needed.
The new data center will be at full capacity in six years, greatly expanding OpenAI’s capacity. (Picture: generated)The operation will be owned by SoftBank’s SB Energy and backstopped by Nvidia guarantees of $105 billion, with the first 800 megawatts of capacity coming online in 2028.
The completed capacity by 2032 would quadruple the compute currently held by OpenAI, said to be about 1.9 GW in 2025, yet growing exponentially.
The facility, built on land previously used for uranium enrichment by the US government, will exclusively use chips and networking from Nvidia, who expects it to represent 1.5 million of their GPUs and revenue of $150—$200 billion, not including future upgrades, according to Jensen Huang.
OpenAI says the data center will use its own energy and recycle its water supply, lessening the load on consumer facilities.
It will also supply the local community with 35,000 construction jobs over six years of building, and 2,500 jobs in operations once it’s finished.
In addition, OpenAI are providing $85 million in Codex credits to Ohio college students for the 2026/27 academic year, in addition to a «community grant fund» of some $40 million.
Dario Amodei wants disproportionally tougher regulation on the heavyweights than for early stage startups. (Picture: Shutterstock)Anthropic CEO Dario Amodei posted a lengthy essay on X during the weekend, where he argues for more regulation of frontier AI models, saying that AI is «structurally a technology that tends to concentrate power.»
He lands in favor of redistribution-like regulations that benefit newer startups and «those catching up,» while disadvantaging the more powerful frontier labs.
Open source models should also be vetted once they reach a certain level, as the White House has agreed to. Open source brings «specific risks,» he argues, although he doubts they will solve the distribution or concentration problem.
On AI’s generally negative public image lately, Amodei says this is fundamentally an issue of trust:
— I think that ordinary people don’t trust companies, governments, or the tech industry and always suspect that we are cooking up some new way to screw them over, he writes.
GLM-5.3 democratizes cybersecurity, but its developer can’t guarantee against misuse. (Picture: Shutterstock)The Chinese open weights model surpasses GPT-5.6 Sol and Fable 5, the sanitized Mythos model, on CyberGym — a widely watched benchmark for cybersecurity — as Z.ai says defensive cyber functions should not be «the privilege of the few,» according to The South China Morning Post.
GLM-5.3 scored 84.5% on the benchmark, slightly higher than Fable 5’s 83.8% and GPT-5.6 Sol’s 83.6%.
On ExploitBench, which is more a measure of a model’s offensive ability, the model trailed the frontiers with 54.5% versus 78% for Fable 5 and 76.5% for GPT Sol.
Z.ai, the lab behind the model, says they will release it as open source with open weights in about two weeks, according to Reuters, after first letting «selected security partners» evaluate it in «controlled settings,» Z.ai writes on security.
Gemini’s latest «workhorse model» does well in coding and web dev, but is not on the frontier. (Picture: Google/generated)After a management shakeup and worries over coding capabilities, Google is out with a new, fast mid-tier model that looks quite capable.
Gemini 3.7 Flash specifically addresses coding («strong gains»), web development («generates more functional layouts») and knowledge work («significantly outperforms 3.6 Flash»).
On Google’s selected benchmarks, it beats Claude Sonnet 5 (not to be confused with the more capable Opus 5) more often than not, goes toe to toe with GPT-5.6 Terra, and handily beats the previous generation, sometimes by a lot.
Pricing is also favorable at $0.75/1M input tokens and $3.75/1M out as an «introductory price» lasting out the year, returning to $1.50 in and $7.50 out in 2027.
It is available as a «workhorse model» in Google’s developer tools, the API, enterprise platform and in Spark for Pro and Ultra subscribers in «supported countries,» meaning that the Spark agent isn’t available in Europe.
The new V4 Pro is one of the best open weights models, and comes with a low price. (Picture: generated)The stealthy upgrade to version 0831 yesterday had many excited, and there was an impressive screenshot of benchmarks circulating and creating buzz all day long, but nobody has been able to verify its origin.
What is confirmed is that DeepSeek has indeed updated its V4 Pro model for the first time since April, and just recently posted this message to its website:
«🎉 The official version of DeepSeek-V4-Pro has been released, featuring significantly enhanced agent capabilities and support for the Responses API and Codex integration. It is now fully available across the web, mobile app, and API; we welcome your testing and feedback.»
Grok has gone from beating the last generation frontier to current gen in about a month. (Picture: generated)A little more than a month after releasing Grok 4.5, that competed mainly on price for GPT-5.5-level performance, SpaceXAI is out with a successor that ups the ante.
The new model reaches parity with GPT-5.6 Sol on the Artificial Analysis benchmark of benchmarks, going from ninth to fourth on the leaderboard, and it also beats Sol on several benches on the Grok presentation.
SpaceXAI claims this proves «frontier level performance,» and it does come in striking distance to the latest state-of-the-art model Opus 5. Considering that GPT Astra is due shortly, this might be a short-lived claim, however.
The new Grok was trained on a «wide range of agentic tasks,» not to mention the corpus of real-word sessions from Cursor, with a focus on knowledge work and general coding.
It is supposedly especially good at «turning a broad product idea into a working first version,» which can then be iterated upon. It is also proficient on long-running, complex tasks and creative visual work, SpaceXAI says.
The main selling point of the model is undoubtedly going to be price/performance, starting at $2 per million input tokens and $6 out, it is astoundingly cheap to run. Prices for GPT-5.6 Sol are $5 for input and $30 for output, which works out at more than twice the input price and 5x on output. Grok 4.6 also has a faster version at double the cost.
The model is out now, and can be found on Cursor, Grok Build and in the API.
Coworkers, not agents. SpaceXAI are marketing Grok Bot as an end-to-end productivity tool. (Picture: SpaceXAI)The app launch talks about coworkers and teammates, not agents and handlers, that you can delegate work to. The «bots» then finish their tasks from start to finish with minimal intervention.
The bots/agents have their own computer in the cloud, like Anthropic’s Cowork and OpenAI’s Work and Codex, and can work across your apps, tools, and websites and use apps with no API or MCP access.
You should be able to talk to them like a fellow worker and delegate entire workflows, from processing invoices in Gmail to adding transcripts and follow-ups in a CRM. There are several workflow demos on x.ai.
To start off with the bots, you just let them monitor your workflow, which they save as routines, and have them do it precisely as you would the next time.
SpaceXAI is internally raving about the product, saying the agents function just like coworkers you can hand real work to, that doesn’t need micromanagement, and works 24/7. They are using it for outbound sales, marketing campaigns, office operations, and, of course, bug fixes.
Grok Bot is available in beta today for users of SuperGrok Heavy (at $300/month), Cursor Ultra (at $200/month), and Cursor Teams Premium (at $120 per user/month). It can be downloaded for macOS, iOS, Windows, and Linux.