Not just hackers, but increasingly sophisticated biological, cyber and influence campaigns have tried to use Claude. (Picture: generated)AI is increasingly used as an efficiency multiplier in intelligence work from Ukraine to China and Yemen, Anthropic’s threat report says.
Attempted threat activity on Claude ranges across the gamut, from cyber to influence and surveillance, scams and fraud, biological misuse and conventional weapons development.
The periodic report details what Anthropic caught and mitigated in slightly older models such as Claude Haiku, Sonnet or Opus, whereas Mythos- and Fable-class models have much stronger safeguards out of the box.
One prominent target for classical cyber operations was Ukraine’s drone industry, trying to create malware for use in disruption and surveillance.
AI is entering a crucially important run of development, ever more researchers say. (Picture: generated)
The Safety and Security Committee of the OpenAI Foundation’s board has a new member in Paul Christiano, a long-running researcher on alignment and safety — and he comes with a warning for the world:
— If we build superintelligence without more robust alignment I expect we will permanently lose control of it, he writes on X, and says — If that happens then most people could die.
The warning comes on the heels of rapid growth in recursive self-improvement recently, and the debate over whether computer systems and AI might gain the possibility to build on themselves at some point in the future. OpenAI reckons they’ll have autonmous AI researchers around 2028, but warns it may come sooner.
His appointment and X essay come just a day after Anthropic researcher Jacob Coxon resigned over alignment, saying his company was «gambling with our lives» by running head-first toward self-improving AI, going viral in the process.
OpenAI officially acknowledges that AI is heading for a new era, saying that: «We’ve reached a new chapter in AI capabilities, and that demands a new chapter for AI policy.» They also add that «this makes it increasingly important that these systems are guided by solid safety, security, alignment, and governance.»
Paul Christiano remains hopeful, ending his X essay by saying that «frontier AI developers have a lot of power to unilaterally improve the situation and lay the groundwork for stronger coordination.»
OpenAI is mum on what their internal agent is otherwise capable of. (Picture: OpenAI)It took an internal model «significantly more capable» than Astra a matter of 88 hours with 10,000 concurrent agents to propose a solution to the Navier Stokes problem, that regulates fluid motion dynamics and has been unsolved for over ninety years.
At the same time, NYU professor Tristan Buckmaster and Anthropic researcher Levent Alpöge had been working on the very same problem using Codex, and while OpenAI hails their «concurrent» work, they claim not to have seen it prior to their solution.
The Navier Stokes solution took 2.7 million messages between the agents and «approximately» 130 billion output tokens, costing roughly $10 million, The BBC reports.
This is the price of progress, OpenAI researcher Noam Brown says, noting that it cost a fortune to win IMO gold before anyone with a $20 subscription could do it a few years later.
— This milestone represents substantial work by mathematicians and AI researchers. However, this is not a culmination, but rather a snapshot in time, of progress on AI development, OpenAI says.
AI will soon be able to improve on itself, OpenAI reckons. (Picture: generated)The average researcher at OpenAI now uses 3.1 agent workdays for every human one, and they say they could reach the level of an autonomous researcher by March 2028.
The use of agents has grown exponentially this year from coding agents in «modest amounts» to integrating agents «daily.» The median researcher uses $600 per day in tokens, whereas the most active ones use more than $7,000 — with a growth of some 124X for researchers.
— I have a strong expectation that this speed of progress could be sustained into recursive self-improvement, Jakub Pachocki, Chief Scientist at OpenAI says in a seperate blog post.
Recursive self-improvement means in essence that an AI model can improve on itself, but OpenAI cautions that they have only just reached «research intern»-status.
An «intern» can carry out research under human supervision that would otherwise take a few days for a skilled researcher.
OpenAI believes agents can be of great help on alignment and deep learning, especially as the models themselves grow larger and more capable — and they are already contributing to faster coding and more experiments.
Automated AI research would «bring down the cost of advanced intelligence so that people worldwide can benefit,» OpenAI says, but it could come at the cost of control and transparency if these agents spin without human oversight.
We’re going to look back on today as the time AGI arrived, Brockman says. (Picture: OpenAI)Astra is a «generational leap» forward, OpenAI tells The Verge on a launch conference, and President Greg Brockman said basically that Artificial General Intelligence has arrived:
— For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.
OpenAI says that Astra is the culmination of years of «big bets» on pre-training, reinforcement learning and alignment.
It tops just about every benchmark out there, sometimes by a lot, and saturates ARC-AGI-3 with a 99.9% score, beating the human baseline, which was thought impossible just weeks ago. It also scores 100% on ExploitBench, which measures how well the model can turn vulnerabilities into exploits, and gets 42.4% on ExploitGym, beating the last state of the art, Fable 5.1, with more than ten points.
It is also the first model measuring as «critical» for cybersecurity on OpenAIs Preparedness Framework, and OpenAI has spent a lot of time on alignment, safeguards and testing. It can effectively find and exploit zero-day (previously unknown) vulnerabilities. A model with «less restrictive» safeguards will be made available through the Daybreak program for vetted researchers and software testers.
Astra will be rolling out to most paid users within «a couple days,» and is currently only available to «a limited set of organizations.» It will cost $10 per million input tokens and $50 for outputs.
Especially the Cyber version outpaces the frontiers at a fraction of the cost. (Picture: Google/generated)Less than three weeks after releasing Gemini 3.7 Flash, its successor is already here. 3.8 Flash excels at analytical legal and financial tasks in benchmarks, outperforming the bigger frontier models.
It slots in with 59 points on Artificial Analysis, between GLM-5.3 and DeepSeek V4, and they note it is fairly cheap ($0.75/1M in/$3.75 out), fast and verbose.
3.8 Flash uses more tokens and does more work even on simple queries than 3.7-Flash, so as a stopgap, 3.7 will continue to be available on the Gemini app and API.
Gemini 3.8 Flash Cyber is the star of this release, however, and shows performance slightly beating the biggest frontier models on CyberGym at a much lower cost.
It is currently used in protecting Google’s own code, where it doesn’t just discover vulnerabilities, it also helps to write the patches for the solution.
Gemini 3.8 Flash is available today for most of the Gemini ecosystem to paying subscribers. The Cyber variant is only available to «trusted defenders.»
The Atlas model can take a single image and generate a detailed, photorealistic, minute-long video in 1440p — and you can freeze the frame and generate a new perspective at any time
World Labs calls it a «multimodal autoregressive diffusion transformer,» which is another way of saying that it puts the image or images into an extensive and knowledgeable world model that can «imagine» the parts that are out of the frame and generate an entire world from very little input.
The striking thing about its outputs is how intuitive and realistic they are, bringing single pictures of, say, a cathedral floor together for a walkthrough in the gallery.
It also seems easy to use. Simply upload between one to 12 images, design a path for the camera, and the model does the rest.
This can be useful for walking through works of art for leisure, exploring a path through a landscape of photos, or for robotics training that no longer requires complex, expensive setups to capture a world in a trainable 3D context.
The model will be used to upgrade outputs from their word building Marble app in future iterations, but is not available for public use just yet. It is in «early access with select partners,» and there is an email list to be notified when it becomes available.
Fable 5.1 will be less annoying to talk to, Anthropic says, after 5.0 had too strict safeguards. (Picture: Anthropic)Fable 5.1 «sets a new standard» in performance, Anthropic says, and does take great leaps on the benchmarks, particularly on agentic coding, knowledge work and science.
The new models score as well or better than Fable 5 for lower cost on the lowest setting. While the sticker price is the same ($10 per M input, $50 out), it is much more efficient and gets the work done with 25% fewer tokens, with 45% less for «highly agentic work.»
Anthropic also says it takes fewer shortcuts in its work, leading to higher-quality outputs compared to Fable 5, and that its safeguards have reduced false positives by 60% — meaning that it will be less restrictive.
The model also stands out in science, scoring more than double Fable 5’s result in the Terminal-Bench-Science benchmark, and while Anthropic aren’t making any specific claims, they do say that «AI models will soon make important contributions to scientific discovery,» and that the new models «offer an early glimpse.»
On cybersecurity, Fable 5.1 is good enough to discover vulnerabilities in code and software, but it is not capable of develop exploits for them. For that, you would need access to Mythos 5.1, which is, as before, the same model with looser safeguards that is only available to trusted partners.
More people are seeing ads on ChatGPT and revenue is growing, OpenAI says. (Picture: generated)In less than 200 days since launch, OpenAI’s ads business has reached a billion dollars in «annualized revenue run rate,» which means that they project this annual number from current revenue.
Ads started as a project only available to partners and agencies, which now counts more than 50 partners — and became a self-serve system in May, that now consists of a «material share» of the business.
ChatGPT only shows ads on the Free and Go tiers, which is used by the vast majority of their billion plus weekly users, and only in the current context of the conversation. It is possible to opt in to more general context from all conversations, though likely few people do that. The ads themselves have no access to said chats or context and don’t influence ChatGPT’s replies.
Having reached the billion dollar benchmark, OpenAI are now expanding the self-serve ad network to India, Europe, the Middle East and North Africa — reaching a total of 40 countries.
Advertisers are «increasingly global,» and OpenAI says non-US revenue is a «growing share of revenue» as many advertisers are now reaching consumers in «multiple markets.»
OpenAI had originally projected $2.5 billion in advertising revenue for 2026, Reuters reports. They reached $100 million in annualized run-rate revenue in April this year, within six weeks since launch.
The prestigious university is reconsidering the fundamentals of teaching due to AI. (Picture: Shutterstock)AI use has been so integrated into student life at the Massachusetts Institute of Technology that they are thinking of a complete overhaul of the education experience. They have widely researched the issue since January 2026, resulting in a report called «AI and Education».
— MIT students use AI frequently and pervasively, the findings say, but list different motives and «strongly mixed feelings, from curiosity, creative inspiration, and gratitude to resignation, concern, and anxiety.»
In less than three years of availability of competent AI, that can complete most undergrad assignments and exams, it has driven to a major shift in campus culture, the report says, and it has happened «suddenly and dramatically, creating a clear sense of urgency.»
The proposed new interface standard will not only make it easier to coordinate and control things like factory floors and lab facilities, but it will also infuse AI agents into the process, allowing for scripting — and adding natural language operations and reasoning.
This could potentially be a huge boon for scientific research and real-world labs, where Claude might one day be able to run end-to-end experiments and manipulate instruments, greatly accelerating the process.
Hugging Face had been looking to raise money and was considering a sale lately. (Picture: generated)Nvidia’s revenues were up 106% to $96.22 billion last quarter, while predicting a 70% increase in the next fiscal year, so agreeing to acquire Hugging Face for $12.9 billion might seem like pocket change.
It is, however, one of Nvidia’s largest acquisitions, Reuters notes, after the buyout was first reported by The Information.
Hugging Face is the «face» of the open source AI movement and maintains a repository of almost all available models, making this a significant infrastructure investment. AI labs treat publishing their weights on it as their official release.
Nvidia is of course not new to Hugging Face or open source, having as good as bought the OSS AI lab Poolside earlier this week to build their own models. They also invested in a $235 million funding round for Hugging Face in 2023 that valued it at $4.5 billion, and tried to invest $500 million in 2025 at a $7 billion valuation, according to The Financial Times.
It appears that Nvidia is somewhat hedging their bets and is increasingly investing in open source, as the frontier AI labs are increasingly developing their own chips, Reuters writes. Nvidia remains a significant investor in closed source providers, though.
The new chip will be deployed at scale within the year. (Picture: OpenAI)Using the measure of performance per watt rather than throughput per second, OpenAI ran three open source models, GPT-OSS, DeepSeek R1 and Kimi K2.5 1T, through their new processor on the InferenceX benchmark from SemiAnalysis.
They found that the processor is wicked fast compared with «leading chips,» which means Nvidia’s offerings, specifically the Grace Blackwell 200 which they show getting trounced in some tests.
The benchmark found that Jalapeño delivers 1.9x more tokens per second when measured per watt on peak efficiency, and is 17.8x faster on higher token density, which translates to raw performance in handling requests.
They also found that end-to-end latency was between 1.7x and 3.4x lower depending on the model, meaning the user will spend less time waiting for responses.
These two measures combined are key, as most chips have to make tradeoffs between latency and throughput, and few can be good at both, OpenAI hardware vice president Richard Ho tells The Verge.
Jalapeño was developed in record time — 9 to 16 months — assisted by AI, and OpenAI says their upcoming Astra model is already busy working on the second generation, which they say is «in deep development.»
OpenAI will deploy and «operate Jalapeño at scale» within their compute infrastructure «by the end of the year.»
Nothing chews through inference tasks as fast as Nvidia’s Groq. By far. (Picture: Nvidia/generated)The never-before-seen feat was achieved running Google’s Gemma 4 31B through Artificial Analysis’ standard tests, and is way ahead of anything on the market.
The only comparable score is that of OpenAI’s GPT-Sol running on Cerebras chips, which achieved 750 tokens per second earlier in August. They called this «Ultrafast mode.»
Nvidia further says that it achieved this output score while maintaining hundreds of thousands of context tokens, and that Groq is some 34X faster than today’s quickest hardware in time to generate 5,000 tokens.
A Groq chip pairs 500 MB of high speed SRAM on die, directly next to the chip, that delivers 150 TB/s throughput. They come in racks of 256 chips stacked together for a total of 40 petabytes of memory bandwidth.
While they are great for inference tasks, GPUs will still be the workhorse of AI data centers, as they can handle both training and later inference — but Jensen Huang of Nvidia recommends setting aside 25% of data center space for the new Groq chip racks, according to CNBC.
The new test scores come as Nvidia is announcing that Groq 3 chips are now in production with Samsung and are generally available.
Even a $4 trillion company is not immune from RAMageddon. (Picture: generated)Nvidia has told its largest customers to expect a price increase for its AI systems of more than 15%, Bloomberg reports.
The hikes come amid soaring prices on memory, components and storage as AI buildouts create unprecedented demand in the market.
The systems involved will be based on both Vera Rubin and Grace Blackwell, and prices will depend on both chip and memory configurations. Increases are expected early next year, writes Reuters.
At the same time, Nvidia is pushing harder into making its own AI services, announcing what is «not an acquisition» and «not an acquihire» — before doing both to AI startup Poolside.
Nvidia will be paying $6 billion to the company to license its software, and is concurrently offering jobs to 109 of its staff. On top of that comes a straight-up investment of $1 billion at a $12 billion valuation.