The new ChatGPT desktop app puts all of OpenAI’s eggs in the same basket. (Picture: OpenAI)UPDATE: OpenAI seems to have removed the chat history and put the chat itself in a popup window. If you don’t like this in the new app, and a lot of people don’t, it leaves the ChatGPT Classic app untouched, which can be used for pure chats like the old version. It will still recieve updates and won’t become obsolete.
—
Along with the release of GPT-5.6 today comes a new app that looks unremarkable at first, but is a one-stop-shop for all of OpenAI’s features.
The latest addition is called ChatGPT Work, which is billed as an agent that gathers information from your apps and workflows and turns them into spreadsheets, presentations, docs or PDFs and just about anything you want — even web apps.
It’s getting integrated into a much more interesting, upgraded ChatGPT desktop app that will combine Codex, Work, agentic browsing and the normal chat interface.
GPT-5.6-Luna is competitive with recently launched cheap models, while Sol is more expensive. (Picture: OpenAI)GPT-5.6 was, as promised, released widely to all of OpenAI’s customers today — and will take around 24 hours to fully propagate.
The model reaches state-of-the-art status through many benchmarks, beats Fable and Mythos often, and scores a record on Arc-Agi-2 of 92.5% while being the first to actually solve a puzzle on Arc-Agi 3, measuring how it handles completely unknown situations.
GPT-5.6 comes in three versions; Sol as the flagship, top rated model, Terra as the mid-tier GPT-5.5-level model at half the price, and Luna as the dirt cheap model, likely comparable with GPT-5.4.
The pricing for Sol is at $5 input/$30 output, Terra is at $2.50 input/$25 output, and Luna ticks in at $1 input and $6 output, placing it in a good position to compete with the recently released Grok 4.5 and Muse Spark 1.1.
The Free and Go tiers on ChatGPT only get access to the Terra model, while Plus, Pro, Business and Enterprise can choose between all three, and set effort levels.
Pro and Enterprise users get Sol Pro in addition, which provides «the highest quality results on complex tasks.»
GPT-Live is better at waiting for its turn to speak, and you can interrupt for a better back-and-forth. (Picture: OpenAI)150 million people use their voice to interact with ChatGPT per week, OpenAI says, and many have waited patiently since 2024 for an upgrade.
The new GPT-Live uses a «full duplex architecture» that separates out the underlying language model and continuously processes speech.
This means it can do «mmm’s» and «yeah’s» during a conversation to indicate it is actively listening, but more importantly it can listen and speak at the same time.
OpenAI says GPT-Live makes decisions «many times per second» on whether to speak, listen, pause or interrupt, letting it engage more naturally in conversations — and do live translation.
The new model delegates more complex tasks to «the latest frontier model» (currently GPT-5.5) for reasoning or web search and gets back to you as soon as it is ready, sometimes showing cue cards for at-a-glance responses in the app.
GPT-Live is available today by tapping the Voice-button in the app or on the web for Go, Plus and Pro subscribers, and there is a GPT-Live-1 mini for the free tiers.
OpenAI has announced the general release of GPT-5.6 tomorrow. (Picture: OpenAI)GPT-5.6 was officially unveiled a little over a week ago, but got hampered by a US government intervention for cyber risks and national security, having been shown in benchmarks to beat Anthropic’s Mythos.
At launch, it got released only to some 20 «trusted partners,» and OpenAI promised a general release «in the coming weeks.»
Now those weeks have passed, the government has lifted its restrictions without further explanation, and both OpenAI’s X account and CEO Sam Altman are posting that it will «launch publicly this Thursday.»
This comes after extensive testing by the Center for AI Standards and Innovation within the Department of Commerce, aided by experts from OpenAI, who stayed on call for «potential questions,» Axios reports.
No account has been given as to what specifically caused the delayed launch or what conditions the government had for its release.
Opus 4.8 was released in May, meaning the gap to Chinese models has closed considerably. (Picture: generated)While not matching GPT-5.5 or Opus 4.8 across the board, the GLM-5.2 is the first open model to come within spitting distance (about 1% on some tests) on coding and cybersecurity tasks at open source benchmarks.
That means Chinese AI lab Z.ai is catching up to cyber capabilities considered by some to be too dangerous to release publicly, and is edging closer to Mythos or GPT-5.6.
The concern is that the new model, released on June 16, is open source and open weight with an MIT license — meaning that anyone can adjust its guardrails and play around with it on any computer capable of running it.
That has researchers worried that China is not only catching up, but that the model might find its way into the hands of bad actors — who will be supercharged when looking for hacking targets, causing what they term «bugmaggedon.»
With Mythos and GPT-5.6 being blocked by the US government, security teams might be tempted to turn to these models at a sixth of the cost of the American frontier, especially as they develop further, notes benchmark provider Semgrep.
The flagship GPT-5.6 Sol is the new state of the art for cyber capabilities. (Picture: Adobe)Caught in yet another government AI debacle, the new models won’t be publicly released until «the coming weeks.» It is getting presented today, and released to «a small group of trusted parties,» said by Axios to number in the twenties.
— We don’t believe this kind of government access process should become the long-term default, writes OpenAI, and says — we [are working] with the Administration to develop the cyber Executive Order framework and a repeatable process for future model releases.
In its presentation, OpenAI details three models in the GPT-5.6 series. Sol is the state-of-the-art Mythos-beating model, while Terra is «a balanced model for everyday work,» and delivers on par with GPT-5.5 at half the cost, and Luna is the cost-focused model «delivering stronger capability at our lowest cost.»
Strong on cyber
The models, particularly Sol, are very strong on cyber defense, and are able to identify security vulnerabilities, but not able to execute automated attacks due to strong safeguards, built using 700,000 A100-equivalent GPU hours of red-teaming.
Sol beats Mythos 5 on coding and cybersecurity and uses fewer tokens for better performance on biology tasks. It also comes with a new max mode, which gives more time for better reasoning, and an ultra mode which uses subagents for more complex work.
Prices for Sol are at $5 input/$30 output, for Terra is $2.50 input/$15 output, and Luna is at $1 input/$6 output, and general availability is up to higher powers.
GPT-5.6 is officially under government review due to «security concerns.» (Picture: generated)First reported by The Information, OpenAI has agreed to the Trump administration’s demands that it stagger the release of the upcoming GPT-5.6 over security concerns.
The model is said to be on par with Anthropic’s Mythos and Fable 5, which were abruptly pulled from the market after an export ban from the administration just two weeks ago.
Axios is quoting sources familiar saying that GPT-5.6 will only be released to a «small set of government-approved partners,» as the White House previews its abilities.
This preview period is expected to last «a couple of weeks,» CEO Sam Altman told staff in an internal OpenAI meeting, TechCrunch reports.
During this time, the administration itself will supposedly approve access to the model on a case-by-case basis, according to The Verge.
The Trump admin issued an executive order on AI about three weeks ago, setting up a voluntary mechanism for AI labs to submit their models for 30 days of pre-release testing. This gave the government 60 days to come up with criteria, but it isn’t quite ready yet — hence this ad-hoc approach.
Lots of talk on socials doesn’t always relate to hard news, but here are some topics making the rounds. (Picture: Adobe)GPT-5.6 in overdrive on the hype machine
The rumor mill is kicking into high gear for OpenAI’s next GPT-5.6. It’s supposed to be largely on par with Fable 5, or kick it to the curb on some tests, according to X leaker Chetaslua who claims to have been testing it for a while.
OpenAI’s chief scientist, Jakub Pachocki, pre-announced the model as a «meaningful improvement,» but would not point to a timeline.
OpenAI’s superapp isn’t just a convenience, it should also be a revenue driver. (Picture: Shutterstock)Much has been said about OpenAI’s coming «superapp,» combining Codex with ChatGPT and now also agents, but the Financial Times is now putting a date certain on it.
The new app is not just a convenience for users who use chat and coding and agents to have it in one place, but it is also a strategic play by OpenAI to funnel them into more lucrative businesses areas — particularly when it comes to agents.
Thibault Sottiaux, who leads OpenAI’s core product and platform efforts, tells the FT that the app will give you a personal agent to do stuff for you, be it personal or work, doing far more than the classic Codex, which now performs computer use and does office tasks.
In fact, 20% of Codex users are now knowledge workers, OpenAI reported last week. This category is also growing three times as fast as other use cases.
The concept is that you will have access to the new app from everywhere you go, from mobile to desktop to the web, so you could even talk to it in your car.
The idea behind the app is to boost revenue and ease enterprise adoption, now counting 2 million Codex users and accounting for 40% of revenue, Reuters reports.
Without touching your Codex box, you can now control it from your phone. (Picture: OpenAI)Instead of lugging your half-open laptop around to stay on top of Codex work, you can now simply leave it in the office and check in from anywhere.
Having Codex on your phone means you can «answer questions, review code, change direction, approve what comes next, or add a new idea,»OpenAI says.
This way, Codex never gets stuck on reviews or permissions, as you will be able to easily guide it from where you are.
You can also chat with it and add specific instructions, like dropping output in a Slack channel or send it on for review by others.
All your credentials, permissions and general setup stays on the machine Codex is running on, while just the updates and requests get routed to your phone.
Codex on ChatGPT for mobile is out today on iOS and Android, available to everyone, but only works with connected macOS machines. A Windows version is «coming soon.»
No more five-page essays when you need a simple answer. (Picture: generated)Replacing GPT-5.3 Instant as the «daily driver» for hundreds of millions of users, the new model touts major changes to common gripes.
For those fed up with three-page essay responses to simple questions, the new model should offer «clearer, more concise answers,» with «stronger» and «tighter» responses that are more «to-the-point.» This should reduce the need for frequent follow-up questions, OpenAI says.
OpenAI says the new Instant will be «more effective» at leveraging history from previous chats, files and Gmail. It should also be showing you its sources when it uses memories.
Hallucinations are supposedly also way down, reducing the rate by 52.5% in internal tests, while inaccurate responses dropped by 37.3%.
The model will become the default option for most ChatGPT users on the Plus and Pro tiers, and is rolling out «over the next two days.» Other plans should get it «soon.»
It took about three weeks for a competing model to hit parity with Mythos. (Picture: Adobe)After a major research paper by the UK’s AI Security Institute found GPT-5.5 a little better than Mythos, Sam Altman moved to limit access to the Cyber version of the model.
The paper probes «vulnerability research and exploitation against realistic targets and modern mitigations» through rigorous tests, and found GPT-5.5 had a pass rate of 71.4%, compared to Mythos’ 68.6% on the most advanced evaluations.
According to the AISI, their test suite proves that Mythos is not a one-off act of brilliance, but part of a wider trend for frontier models. They say «we should expect further increases in cyber capability from models in the near future, potentially in quick succession.»
At the same time, Sam Altman posted on x.com that OpenAI will indeed follow Anthropic’s lead on limiting access to GPT-5.5-Cyber to «critical cyber defenders:»
— We will work with the entire ecosystem and the government to figure out trusted access for cyber; we want to rapidly help secure companies/infrastructure, Altman wrote.
Today, OpenAI is revealing their research on the issue, and can reveal that this was indeed real. Starting with GPT-5.1, the models did definitively prefer using «goblins» in their replies.
The culprit was the «nerdy» personality, which debuted with the launch of the 5.1-model and had increased «goblin»-mentions by 175% and «gremlin» by 52%. And by GPT-5.4, «goblin»-use had balloned by 3,881.4%, causing consternation at OpenAI.
The error seems to stem from rewarding a «playful style» with creature references, and this has since propagated through later releases.
The «nerdy» personality was retired in March after GPT-5.4 was released, but goblins snuck into the training data for GPT-5.5, too — forcing the system prompt to «Never talk about goblins, gremlins, raccoons, trolls, ogres, pigeons, or other animals or creatures unless it is absolutely and unambiguously relevant to the user’s query.»
We can expect more models at a faster pace from OpenAI going forward. (Picture: Adobe)GPT-5.5 came out just six weeks after GPT-5.4, so naturally, people are wondering if this is part of a new trend of quicker releases:
— Yes, we expect quite rapid continued progress, says OpenAI Chief Scientist Jakub Pachocki on a call related to today’s release, according to Tae Kim on Substack.
— We see pretty significant improvements in the short term, extremely significant improvements in the medium term, he adds.
Sam Altman has also weighed in, saying that «We believe in iterative deployment; although GPT-5.5 is already a smart model, we expect rapid improvements.»
Altman goes on to say that this is also part of a safety strategy, arguing that iterative releases make it easier for the world to prepare for advances in AI.
From chatbot to «research partner,» ChatGPT-5.5 steps up with complex reasoning. (Picture: Adobe)Some advanced models lose speed when upgraded, but not so for OpenAI’s latest. GPT-5.5 handles more demanding tasks at a faster clip — and «excels» at agentic coding, computer use, knowledge work and scientific research.
It also matches the GPT-5.4 on token latency, «while performing at a much higher level of intelligence,» OpenAI says.
Codex is much improved by the new model, with better code debugging, document and spreadsheet creation, use of software and moving across toolsets.
It’s also strong on benchmarks, using significantly fewer tokens to complete the same Codex tasks as GPT-5.4.
On computer use and scientific research, it should get higher quality results with fewer tokens or retries, and requires less guidance, OpenAI says.
The model is more expensive than GPT-5.4, clocking in at $5 per million input tokens and $30 per million output tokens — almost twice the cost.