Anthropic previews model interface for AI control of physical devices

The proposed new interface standard will not only make it easier to coordinate and control things like factory floors and lab facilities, but it will also infuse AI agents into the process, allowing for scripting — and adding natural language operations and reasoning.

This could potentially be a huge boon for scientific research and real-world labs, where Claude might one day be able to run end-to-end experiments and manipulate instruments, greatly accelerating the process.

Continue reading “Anthropic previews model interface for AI control of physical devices”

OpenAI’s Jalapeño beats Nvidia’s Blackwell in performance per watt

The new chip will be deployed at scale within the year. (Picture: OpenAI)
Using the measure of performance per watt rather than throughput per second, OpenAI ran three open source models, GPT-OSS, DeepSeek R1 and Kimi K2.5 1T, through their new processor on the InferenceX benchmark from SemiAnalysis.

They found that the processor is wicked fast compared with «leading chips,» which means Nvidia’s offerings, specifically the Grace Blackwell 200 which they show getting trounced in some tests.

The benchmark found that Jalapeño delivers 1.9x more tokens per second when measured per watt on peak efficiency, and is 17.8x faster on higher token density, which translates to raw performance in handling requests.

They also found that end-to-end latency was between 1.7x and 3.4x lower depending on the model, meaning the user will spend less time waiting for responses.

These two measures combined are key, as most chips have to make tradeoffs between latency and throughput, and few can be good at both, OpenAI hardware vice president Richard Ho tells The Verge.

Jalapeño was developed in record time — 9 to 16 months — assisted by AI, and OpenAI says their upcoming Astra model is already busy working on the second generation, which they say is «in deep development.»

OpenAI will deploy and «operate Jalapeño at scale» within their compute infrastructure «by the end of the year.»

Read more: OpenAI’s report, X post. The Verge, Bloomberg (paywalled) and The Register. Discussion on Hacker News and r/Singularity.

Nvidia’s inference rack Groq 3 LPX delivers record 3,431 tokens per second

Nothing chews through inference tasks as fast as Nvidia’s Groq. By far. (Picture: Nvidia/generated)
The never-before-seen feat was achieved running Google’s Gemma 4 31B through Artificial Analysis’ standard tests, and is way ahead of anything on the market.

The only comparable score is that of OpenAI’s GPT-Sol running on Cerebras chips, which achieved 750 tokens per second earlier in August. They called this «Ultrafast mode.»

For more normal hardware setups, Opus 5 gets 58.8 tokens per second, and GPT-5.6-Sol clocks in at 74.4 per second, while the «faster» Gemini 3.7 Flash gets 371.1 throughput tokens.

Nvidia further says that it achieved this output score while maintaining hundreds of thousands of context tokens, and that Groq is some 34X faster than today’s quickest hardware in time to generate 5,000 tokens.

A Groq chip pairs 500 MB of high speed SRAM on die, directly next to the chip, that delivers 150 TB/s throughput. They come in racks of 256 chips stacked together for a total of 40 petabytes of memory bandwidth.

While they are great for inference tasks, GPUs will still be the workhorse of AI data centers, as they can handle both training and later inference — but Jensen Huang of Nvidia recommends setting aside 25% of data center space for the new Groq chip racks, according to CNBC.

The new test scores come as Nvidia is announcing that Groq 3 chips are now in production with Samsung and are generally available.

Read more: Nvidia’s announcement, production note. Writeups on CNBC and The Register.

OpenAI’s next hardware device is a screenless, smart home speaker

The new device is supposedly «something to take a bite of» per previous reporting, and will look nothing like this. (Picture: generated)
Jony Ive’s first hardware device for OpenAI will be a movable, screen-free speaker with a camera and sensors to understand the environment, Bloomberg reports. It can control other smart home appliances and lets you talk to a customized, more personalized ChatGPT.

The speaker is slated for launch later this year, but won’t actually ship until sometime in 2027. It is one of about five devices slated for launch soon, Bloomberg says.

It will proactively reach out once it gets to know you, comes with a battery to move easily between rooms, and can do things like play music, answer messages or stay chatting, apparently with «personality.»

It will be powered by an advanced version of GPT-Live, and will be more than just talking to ChatGPT. Bloomberg says it will be more of a human-like companion, becoming an «expert on the user» over time.

The speaker also has mechanical elements that move on their own to «connect on a humanlike level with users»

OpenAI believes the device won’t be affected by the recent lawsuit on trade secrets from Apple, Bloomberg says, as it is a new class of devices not found anywhere else. Apple does make speakers, though, like the HomePod.

Read more: Bloomberg (paywalled), Slashdot, MacRumors, TechCrunch, and The Verge.

Apple sues OpenAI over hardware secrets, calls unit «rotten to the core»

A jury will have to decide if OpenAI is guilty of systematic theft of Apple secrets (Picture: generated)
The lawsuit from Apple is a damning indictment of OpenAI for stealing technology, prototypes, business processes and even partner manufacturers through Apple’s former employees, 9to5Mac reports.

At the heart of it lies OpenAI’s chief hardware officer Tang Tan, who left his role as VP of Product Design at Apple in 2024 to work with Jony Ive, later joining OpenAI in the Io acquisition, and Chang Liu, who spent eight years at Apple as a Senior System Electrical Engineer before leaving for OpenAI in January.

The pair is accused of accessing confidential Apple systems and downloading «detailed information about unreleased products, engineering presentations, technical specifications, and proprietary project data,» The Verge reports.

The two have also been holding some interesting job interviews, Apple claims, where hires from Apple have been encouraged to lay out systems planning and to log into confidential systems and share trade secrets and restricted paperwork.

There are 400 ex-Apple employees at OpenAI, and the whole hardware business is «rotten to its core by its illegal reliance on misappropriated trade secrets,» the lawsuit alleges.

«We have no interest in other companies’ trade secrets,» OpenAI spokesperson Drew Pusateri tells The Guardian.

Read more: The actual filing. More on 9to5Mac, The Verge, The Guardian, TechCrunch and Daring Fireball. Discussion on r/Singularity and Hacker News.

Apple says security updates must be more frequent due to AI hacking

AI accelerated hacking is changing how Apple deals with system patches. (Picture: generated)
After releasing a 26.5.2 update across its devices to fix over 25 bugs today, Apple said the fixes were originally intended for the next point release of their operating systems — the 26.6.

Concerns about AI-accelerated hacking tools made the company push up the updates sooner in a dedicated release, according to Reuters.

This is the new reality, they say, where the time to develop malware has been greatly reduced by AI — and the time window for fixing bugs has likewise decreased.

Apple did not say if any of the vulnerabilities they rectified in the latest release had been exploited in the wild yet, which used to be the criteria for a rapid response.

The company is one of the «trusted partners» on Project Glasswing, and enjoys access to Anthropic’s Mythos Preview model to probe and detect flaws in their software, MacRumors notes, but it is not known if Mythos was used in this particular case.

Read more: Reuters, MacRumors, 9to5Mac.

OpenAI’s new hardware device is a shortcut keypad for Codex

The Codex keypad looks like an iteration of an existing product, not something entirely new. (Picture: OpenAI)
The OpenAI Developers account on X teased a video of a square device with backlit keys made in collaboration with Work Louder, along with the date July 15th.

It’s likely a small keypad for common Codex shortcuts, as mentioned by the post, and it will be the first hardware device (co-)developed by OpenAI. Not a flashy new phone or the highly anticipated AI hardware device, which would take years to develop.

Work Louder is mostly famous for mechanical keyboards and custom keypads, and makes a $144 13-key wireless pad, complete with a tiny joystick.

The device being teased in the video looks like a carbon copy of the Creator Micro 2 pad described above, save for the backlit keys.

That should offer a hint for how OpenAI will be playing its hardware — as a premium Codex add-on, likely also with a premium price.

Read more: The X post, The Verge, 9to5Mac.

OpenAI delivers Jalapeño, a state-of-the-art chip made in only nine months

The Jalapeño chip offers substantially better performance per watt than anything on the market. (Picture: OpenAI)
The new inference chip, developed with help from OpenAI’s models in collaboration with Broadcom, has compressed what is normally a glacial, years-long process into just a short sprint.

— We believe [this] to be the fastest ASIC development cycle ever achieved in high-performance advanced semiconductors, OpenAI says.

The chip is designed specifically for OpenAI’s needs on current and upcoming language models, and to combine the power of current AI accelerators with the shorter latency of «specialized systems,» like Cerebras’ hardware.

— The world is moving to a compute-powered economy, says Greg Brockman, President and Co-Founder of OpenAI. — Jalapeño is part of our long-term full-stack infrastructure strategy to make compute more abundant.

The new silicon will be deployed at gigawatt-scale data centers just as fast as it was made, starting with Microsoft «and other partners» already in 2026, Broadcom says. It is only the first design in what the companies expect to be generations of products.

Jalapeño is already running LLM workloads in the lab, delivering «substantially better» performance per watt than «the current state-of-the-art,» OpenAI says.

Read more: OpenAI’s presentation, CNBC, CNN, Reuters, and TechCrunch.

Nvidia launches Arm-based RTX Spark SoC for Windows-based AI computers

Nvidia promises a new paradigm in how we use computers, but offers little detail. (Picture: Nvidia)
Claiming a new era in PC processing power, the new system chip is custom-built for AI workflows and is «a New Beginning for Personal Computers,» Nvidia says.

According to Nvidia, people with the RTX Spark installed can simply talk to their computer to get stuff done. A designer can get AI to evolve their sketch all the way to a finished 3D model and a movie with Adobe tools with agents doing all the lifting, for example, The Verge reports.

Continue reading “Nvidia launches Arm-based RTX Spark SoC for Windows-based AI computers”

OpenAI is building a «fundamentally changing» agent-first phone, Kuo says

This is how Ming-Chi Kuo imagines a new agent-first interface. (Picture: Ming-Chi Kuo)
Supply chain analyst Ming-Chi Kuo, mostly known for breaking early Apple news, has published an article on x.com claiming that OpenAI is already working with MediaTek and Qualcomm to develop a phone with an agentic interface.

— Smartphones will remain the largest-scale device category for the foreseeable future, he writes, and shipments of high-end phones are around 300-400 million units a year — a mass market OpenAI would be keen on reaching.

The phone itself should be an agent-first experience, «fundamentally changing how people think of smartphones,» Kuo says.

This means moving away from apps to agents, from icons to tasks, and then doing away with the well-worn grid interface in favor of an agent-powered stream layout.

Along with the chip giants on board, systems integrator Luxshare is slated as their main manufacturing partner, aiming for a full system spec late this year or early 2027. Mass production is «expected» in 2028, meaning the project is moving fast.

Read the full story on x.com.

Arm releases its first physical silicon chip, the AGI CPU, for agentic inference

Arm says agent workflows are set to rise four times, and their new CPU is tailor made for the process. (Picture: Arm)
Precisely catching the fastest growing trend in AI computing, Arm says its new CPU is tailor made for agent workloads.

Co-developed with Meta, the chip is claimed to deliver twice the performance per rack compared to x86 platforms.

Agentic AI compute is expected to require more than four times the current capacity per gigawatt in data centers, and both Arm and Meta expect the design to iterate across several generations.

The AGI CPU is projected to lift Arm’s revenue by «billions» of dollars, Reuters reports, and has over fifty launch partners, including OpenAI, Amazon AWS, Google Cloud, and Meta.

Read more: Arm presser, Meta presser, product page, Reuters, and The Verge.

Nvidia will sell $1 trillion of its AI chips by 2027, launches inference rack

The Groq 3 LPU has only 500 MB of memory, but it’s SRAM flying at 150 TB/s. (Picture: Nvidia)
The Blackwell and Rubin series of chips are selling like hotcakes, the Nvidia CEO says at the Games Developer Conference, as he doubles the previous guidance of $500 billion in sales and justifies a market valuation topping $4 trillion.

Huang’s most interesting offering at the show was the new Groq 3 LPX, a custom rack made for inference loads.

Continue reading “Nvidia will sell $1 trillion of its AI chips by 2027, launches inference rack”

Meta announces new processor generation focused on inference

Meta’s MTIA-chips are supposed to support recommendations and AI for «billions» of people. (Picture: Meta)
Tapping their long-time partner Broadcom, Meta’s new in-house chips are built to scale from recommendation engines to advanced AI workflows.

The new generation Meta Training and Inference Accelerator (MTIA) isn’t built to replace the chips sourced from AMD or Nvidia, but are intended to supplement them and achieve «the lowest possible price.»

MTIA-chips are taking a different tack on developing AI models, optimizing for inference — the process of answering AI queries — instead of training. Most chips are customized for training, which is more compute intensive, but the majority of a chip’s life is spent putting together answers in production.

— Chip designs are based on projected workloads, but by the time the hardware reaches production — often two years later — those workloads may have shifted substantially, Meta says in their press release.

The new chips «have either already been deployed or are scheduled for deployment in 2026 or 2027,» Meta says — and they don’t disclose just how many of these they are making.

Read more: Meta’s press release. Writeups on Reuters and CNBC.

OpenAI’s first device will reportedly be a pocket-sized AI speaker, due in 2027

Does the world really need another smart speaker? Ive and Altman certainly think so. (Picture: screenshot)
After buying former Chief Apple Designer Jony Ive’s Io design lab in May, 2025, Altman and Ive have been teasing a breakthrough hardware device said to be something to take a bite of — and speculation has abounded.

Now The Information (paywalled) is citing sources from an all hands meeting at OpenAI touting an early prototype of a smart speaker.

Continue reading “OpenAI’s first device will reportedly be a pocket-sized AI speaker, due in 2027”

ByteDance is developing in-house AI chips, to be manufactured by Samsung

Nvidia chips are available in China, but users need permission to buy them. (Picture: Adobe)
Not much is known about the AI inference chips, or how they compare to Nvidia’s offerings, but ByteDance is going to be making about 100,000 of them «this year,» and then scale up to 350,000 units, according to Reuters.

ByteDance has been known to work with US chip producer Broadcom, and started seriously hiring chip specialists in 2022.

The new chips are set to be produced with Samsung in a deal that includes memory chips, which definitely sweetens the deal.

Production is advanced enough that Reuters’ sources say engineering samples are due by late March, which is the last stage before production.

A spokesperson for the company does not deny the report outright, but claims the information is «inaccurate,» Reuters writes.

Most US frontier labs are developing their own chips, as is Alibaba and Baidu.

Read the scoop at Reuters.