AI · AI at work
ChatGPT Work: What OpenAI's Work Agent Actually Delivers
ChatGPT Work turns the chat window into an agent that finishes the job – powered by the new GPT-5.6 models and a voice mode that listens and speaks at once. What the update delivers, what it consumes and what actually reaches Europe. As of July 2026.
By Boaz Lichtenstein Prefer us on Google

For two years the promise of AI assistants was: ask me something. Since 9 July 2026, OpenAI’s promise reads: get it done. ChatGPT Work is an agentic mode that takes a goal, breaks it into steps, gathers context from connected apps and files, and delivers a finished result – a document, a spreadsheet, a deck, a report. Underneath sits the new GPT-5.6 model family, and since 23 July the whole thing can be directed by voice. This piece sorts out what has substance, what it consumes and what genuinely reaches Europe. As of July 2026.
Key takeaways
- Work is not a new subscription tier but a mode within the paid plans except Free and Go – Pro, Pro Lite, Enterprise and Edu first, Plus and Business days later.
- Underneath runs GPT-5.6 in three tiers: Sol (flagship), Terra (everyday), Luna (fast and cheap) – the number marks the generation, the name the capability class.
- The new GPT-Live voice models listen and speak simultaneously instead of running the old chain of transcription, model and speech output – available in Work and Codex on the desktop since 23 July. A public API is not out yet.
- The real leverage lies in connectors and scheduled tasks: Slack, Drive, inbox, CRM, repos – plus runs that start on a schedule or on an event.
- For Europe: Work yes, Sites no (blocked in the EEA, Switzerland and the UK); data and inference residency exist only for Enterprise and Edu – and only through sales.
What ChatGPT Work is – and what it is not
The distinction is worth drawing, because OpenAI has shipped three similar-sounding things in twelve months. Projects is a folder: a container for chats, files and instructions around one endeavour. Deep Research is a research brief that ends in a report. Codex is the coding agent for repositories, diffs and tests.
Work is the mode where all of that converges and aims at an artefact. You describe an outcome – “prepare the quarterly report for sales, figures from the CRM, comments from the Slack channel, same format as last time” – and the agent plans on its own, fetches the sources and stays with the job for hours if it is a big one. It runs in the cloud; you can close the window.
The desktop app adds what a cloud cannot: reading and writing local files, driving a built-in browser and operating the computer itself. Those three capabilities are why the desktop client now earns its place for Work users – and why Enterprise and Edu workspaces got a two-week run-up during which Work stayed switched off by default.
What Work is not: a replacement for specialist software, and no reason to pass results on unchecked. The agent delivers a complete first version – not the final sign-off. How to anchor that distinction organisationally is covered in the practical guide to AI adoption.
Sol, Terra, Luna: the three engines underneath
GPT-5.6 became generally available the same day – after two weeks in which access was limited to roughly two dozen organisations vetted by the US agency CAISI. What is notable is less the model than the structure: instead of one version there are three tiers, meant to advance on their own schedules from here on.
| Tier | Intended for | API price per 1M tokens (in/out) |
|---|---|---|
| Sol | Flagship: heavy reasoning, large codebases, long runs | $5 / $30 |
| Terra | Balanced everyday work – the sensible default | $2.50 / $15 |
| Luna | Fast and cheap: volume, classification, summarisation | $1 / $6 |
The naming logic is more practical than it sounds: the number marks the generation, the celestial body the capability class. A leap in one tier no longer forces the whole family into a rename.
OpenAI states that Sol is 54 per cent more token-efficient than its predecessor on agentic coding tasks. That is a vendor claim, independently unverified – but it is the right metric: for an agent that pushes a task through hundreds of steps, the bill is decided less by the price per token than by how many tokens the same outcome takes. Where GPT-5.6 stands among the frontier models is covered in Fable 5 vs. Kimi K3.
Voice: what “listening and speaking at once” changes
A day before Work, on 8 July, came GPT-Live-1 and GPT-Live-1 mini. The difference from the previous voice mode is architectural, not cosmetic: Advanced Voice was a chain of three systems – speech to text, text to model, answer back to speech. Every link cost time, and none of them could listen while speaking. Hence the familiar symptoms: interrupting barely worked, pauses felt mechanical, a quick “hang on” got lost.
GPT-Live is full duplex – one system that listens and speaks while doing both. The mini model has replaced Advanced Voice as the default; the larger model sits on the paid plans. Since 23 July voice also works in Work and Codex in the desktop app: you can interject into a running agent, sharpen the brief or direct several parallel runs without switching focus.
Two honest caveats. First: in the live translation demos the quality was
audibly uneven – the Hindi output drew criticism for a heavy American
accent and a wooden delivery. Independent tests for other languages are
not in yet. Second: there is no public API for GPT-Live, though
OpenAI has announced one. Anyone building voice applications today still
uses the Realtime API with gpt-realtime-2.1 – $32 per million audio
input tokens and $64 per million audio output tokens, or $10 and $20 on
the mini variant. As a rough figure, that is around five cents per
conversation minute on the larger model and just under two on the
smaller one.
Connectors and scheduled tasks: where the value appears
The model tiers are the headline; the value sits elsewhere. Work draws
its context from a plugin directory covering the usual work systems:
Slack and Teams, Google Drive and SharePoint, Gmail and Outlook,
Salesforce, GitHub, Linear, Figma, Zoom, Dropbox. Individual sources can
be addressed with @ instead of leaving the agent to guess where to
look.
On top of that come scheduled tasks: jobs that run once, on a recurring schedule or on an event – and that can watch sources over time. That shifts the character of the tool. No longer “let me ask something”, but:
- A Monday-morning sales report built from CRM figures and channel comments, in last week’s format.
- A briefing before every client meeting, assembled from the email thread, the current quote and recent tickets.
- A watch that fires when defined pages or a repository change.
The limiting factor is almost never model quality but how precisely the agent reaches the right context – the subject spelled out in context engineering. Anyone who has not tidied their system landscape will mostly get more confident errors from a stronger model. The stages before this one – from chatbot to acting system – are described in AI agents at work.
What it costs – and why the maths looks different
There is no surcharge and no new tier. What there is, is a consumption pattern unlike chat:
| Plan | ChatGPT Work | Note |
|---|---|---|
| Free, Go | no | The new desktop app yes, the Work mode no |
| Plus, Business | yes | Followed a few days after the other plans |
| Pro, Pro Lite | yes | Included from day one |
| Enterprise, Edu | yes | Two-week preview, switched off by default at first |
Billing runs through the same agent-work allowance that Codex uses. A job that runs for an hour, reads dozens of files and discards several drafts eats into that allowance noticeably faster than a hundred chat messages – and no per-task rates have been published. Workspaces get spend controls per group and per user; everyone else is left with: run one realistic job, then extrapolate. Why context length and repetition are the real cost drivers is explained in understanding AI costs.
The European view: processing agreements, residency, reality
Work itself is generally available, Europe included. It gets interesting one layer down.
Contractually: as soon as personal data is involved, you need a data processing agreement. That exists for Business, Enterprise and the API – not for a private Plus account. An agent with access to a company inbox running on a personal account is not a grey area; it is simply the wrong contractual basis.
Technically: data residency in Europe has existed since 2025, and since January 2026 there is inference residency as well – processing then happens in European data centres too, not just storage. Both are aimed at Enterprise and Edu, are configured per project and go through sales rather than a toggle in your account. If you need it, you have to negotiate for it.
And the gap: Sites, the website builder OpenAI markets as part of Work, is unavailable in the EEA, Switzerland and the UK, and supports neither data nor inference residency. What that means for Europe is covered in ChatGPT Sites.
The wider point is the same as with any cloud service: an agent allowed to read inbox, storage and chat sees more than any single tool before it. Which content to entrust to someone else’s infrastructure at all is sorted out in ChatGPT and privacy.
Governance: who may do what – and who answers for the agent
OpenAI ships more tooling for workspaces than the headlines suggest: restrictions on plugins, browser and network access, an automatic review layer in which a strong model checks risky actions, a compliance interface for IT, and spend limits.
None of that replaces a rule of your own, and the most important one is uncomfortable: an agent that reads external content and is allowed to act is attackable through exactly that content. A doctored email, a manipulated document, a page with hidden instructions – the attack path is called prompt injection, it is architectural, and no vendor has solved it completely; the mechanics and the countermeasures that actually work are in prompt injection: securing AI agents properly.
Three guardrails cover most of the practice:
- Read broadly, write narrowly. Generous access to sources, minimal and explicit rights to change and send.
- Human sign-off for anything outbound. Emails, bookings, commits, publications. No exceptions, however tedious.
- Log it and spot-check it. What an agent did has to be reconstructable – otherwise the liability question has no answer.
That adoption fails on people rather than technology is the rule, not the exception – see AI acceptance in teams.
Work, Cowork, Copilot: where OpenAI sits in the field
Work is the answer to two products that arrived earlier. The three differ less in intelligence than in where they do their work:
| Approach | Runs | Strength | Available since |
|---|---|---|---|
| ChatGPT Work | in the cloud, app-centric | Breadth of connectors and output formats | 9 July 2026 |
| Claude Cowork | on the desktop, on the local file system | Long documents, work in real folders | January 2026 (Mac), February (Windows) |
| Copilot Cowork | inside the Microsoft 365 tenant | Existing Microsoft stack and permissions | generally available since 16 June 2026 |
Crowning a winner would be dishonest, because the decision rarely turns on the model. It turns on where the data lives that the agent is meant to work with: in the Microsoft tenant, on local drives, or spread across a dozen cloud services. The third pattern is the case Work was built for.
Who Work pays off for today
- Teams with many systems and recurring reports. This is exactly where connectors plus scheduled tasks earn their keep.
- Solo operators with a research and documentation load. An agent that turns scattered sources into a presentable first draft saves real hours.
- Organisations already standardised on ChatGPT. The mode costs no additional contract, only allowance.
Anyone whose access permissions are not under control, who is unwilling to define approval steps for outbound actions, or who strictly needs European inference residency without negotiating an Enterprise agreement, should postpone adoption rather than force it.
The bottom line
ChatGPT Work is the most coherent version yet of what OpenAI has been signalling for two years: the assistant stops answering and starts delivering. The three model tiers make costs more predictable, voice finally removes the stop-start rhythm from the interaction – and the connectors are the part that decides whether any of it is useful.
The sober half: consumption is hard to predict, the voice API is still missing, and the European fine print – contract, residency, the blocked Sites feature – has to be actively clarified rather than assumed. If you start today, start with one recurring job, narrow write permissions and a person who reads the output. Exactly as the practical guide to AI adoption recommends for every stage before this one.