Skip to content
Diese Seite gibt es auch auf Deutsch.Zur deutschen Version

AI · AI at work

ChatGPT Work: What OpenAI's Work Agent Actually Delivers

ChatGPT Work turns the chat window into an agent that finishes the job – powered by the new GPT-5.6 models and a voice mode that listens and speaks at once. What the update delivers, what it consumes and what actually reaches Europe. As of July 2026.

By Boaz Lichtenstein Prefer us on Google

Article image: ChatGPT Work: What OpenAI's Work Agent Actually Delivers

For two years the promise of AI assistants was: ask me something. Since 9 July 2026, OpenAI’s promise reads: get it done. ChatGPT Work is an agentic mode that takes a goal, breaks it into steps, gathers context from connected apps and files, and delivers a finished result – a document, a spreadsheet, a deck, a report. Underneath sits the new GPT-5.6 model family, and since 23 July the whole thing can be directed by voice. This piece sorts out what has substance, what it consumes and what genuinely reaches Europe. As of July 2026.

Key takeaways

  • Work is not a new subscription tier but a mode within the paid plans except Free and Go – Pro, Pro Lite, Enterprise and Edu first, Plus and Business days later.
  • Underneath runs GPT-5.6 in three tiers: Sol (flagship), Terra (everyday), Luna (fast and cheap) – the number marks the generation, the name the capability class.
  • The new GPT-Live voice models listen and speak simultaneously instead of running the old chain of transcription, model and speech output – available in Work and Codex on the desktop since 23 July. A public API is not out yet.
  • The real leverage lies in connectors and scheduled tasks: Slack, Drive, inbox, CRM, repos – plus runs that start on a schedule or on an event.
  • For Europe: Work yes, Sites no (blocked in the EEA, Switzerland and the UK); data and inference residency exist only for Enterprise and Edu – and only through sales.

What ChatGPT Work is – and what it is not

The distinction is worth drawing, because OpenAI has shipped three similar-sounding things in twelve months. Projects is a folder: a container for chats, files and instructions around one endeavour. Deep Research is a research brief that ends in a report. Codex is the coding agent for repositories, diffs and tests.

Work is the mode where all of that converges and aims at an artefact. You describe an outcome – “prepare the quarterly report for sales, figures from the CRM, comments from the Slack channel, same format as last time” – and the agent plans on its own, fetches the sources and stays with the job for hours if it is a big one. It runs in the cloud; you can close the window.

The desktop app adds what a cloud cannot: reading and writing local files, driving a built-in browser and operating the computer itself. Those three capabilities are why the desktop client now earns its place for Work users – and why Enterprise and Edu workspaces got a two-week run-up during which Work stayed switched off by default.

What Work is not: a replacement for specialist software, and no reason to pass results on unchecked. The agent delivers a complete first version – not the final sign-off. How to anchor that distinction organisationally is covered in the practical guide to AI adoption.

Sol, Terra, Luna: the three engines underneath

GPT-5.6 became generally available the same day – after two weeks in which access was limited to roughly two dozen organisations vetted by the US agency CAISI. What is notable is less the model than the structure: instead of one version there are three tiers, meant to advance on their own schedules from here on.

Tier Intended for API price per 1M tokens (in/out)
Sol Flagship: heavy reasoning, large codebases, long runs $5 / $30
Terra Balanced everyday work – the sensible default $2.50 / $15
Luna Fast and cheap: volume, classification, summarisation $1 / $6

The naming logic is more practical than it sounds: the number marks the generation, the celestial body the capability class. A leap in one tier no longer forces the whole family into a rename.

OpenAI states that Sol is 54 per cent more token-efficient than its predecessor on agentic coding tasks. That is a vendor claim, independently unverified – but it is the right metric: for an agent that pushes a task through hundreds of steps, the bill is decided less by the price per token than by how many tokens the same outcome takes. Where GPT-5.6 stands among the frontier models is covered in Fable 5 vs. Kimi K3.

Voice: what “listening and speaking at once” changes

A day before Work, on 8 July, came GPT-Live-1 and GPT-Live-1 mini. The difference from the previous voice mode is architectural, not cosmetic: Advanced Voice was a chain of three systems – speech to text, text to model, answer back to speech. Every link cost time, and none of them could listen while speaking. Hence the familiar symptoms: interrupting barely worked, pauses felt mechanical, a quick “hang on” got lost.

GPT-Live is full duplex – one system that listens and speaks while doing both. The mini model has replaced Advanced Voice as the default; the larger model sits on the paid plans. Since 23 July voice also works in Work and Codex in the desktop app: you can interject into a running agent, sharpen the brief or direct several parallel runs without switching focus.

Two honest caveats. First: in the live translation demos the quality was audibly uneven – the Hindi output drew criticism for a heavy American accent and a wooden delivery. Independent tests for other languages are not in yet. Second: there is no public API for GPT-Live, though OpenAI has announced one. Anyone building voice applications today still uses the Realtime API with gpt-realtime-2.1 – $32 per million audio input tokens and $64 per million audio output tokens, or $10 and $20 on the mini variant. As a rough figure, that is around five cents per conversation minute on the larger model and just under two on the smaller one.

Connectors and scheduled tasks: where the value appears

The model tiers are the headline; the value sits elsewhere. Work draws its context from a plugin directory covering the usual work systems: Slack and Teams, Google Drive and SharePoint, Gmail and Outlook, Salesforce, GitHub, Linear, Figma, Zoom, Dropbox. Individual sources can be addressed with @ instead of leaving the agent to guess where to look.

On top of that come scheduled tasks: jobs that run once, on a recurring schedule or on an event – and that can watch sources over time. That shifts the character of the tool. No longer “let me ask something”, but:

  • A Monday-morning sales report built from CRM figures and channel comments, in last week’s format.
  • A briefing before every client meeting, assembled from the email thread, the current quote and recent tickets.
  • A watch that fires when defined pages or a repository change.

The limiting factor is almost never model quality but how precisely the agent reaches the right context – the subject spelled out in context engineering. Anyone who has not tidied their system landscape will mostly get more confident errors from a stronger model. The stages before this one – from chatbot to acting system – are described in AI agents at work.

What it costs – and why the maths looks different

There is no surcharge and no new tier. What there is, is a consumption pattern unlike chat:

Plan ChatGPT Work Note
Free, Go no The new desktop app yes, the Work mode no
Plus, Business yes Followed a few days after the other plans
Pro, Pro Lite yes Included from day one
Enterprise, Edu yes Two-week preview, switched off by default at first

Billing runs through the same agent-work allowance that Codex uses. A job that runs for an hour, reads dozens of files and discards several drafts eats into that allowance noticeably faster than a hundred chat messages – and no per-task rates have been published. Workspaces get spend controls per group and per user; everyone else is left with: run one realistic job, then extrapolate. Why context length and repetition are the real cost drivers is explained in understanding AI costs.

The European view: processing agreements, residency, reality

Work itself is generally available, Europe included. It gets interesting one layer down.

Contractually: as soon as personal data is involved, you need a data processing agreement. That exists for Business, Enterprise and the API – not for a private Plus account. An agent with access to a company inbox running on a personal account is not a grey area; it is simply the wrong contractual basis.

Technically: data residency in Europe has existed since 2025, and since January 2026 there is inference residency as well – processing then happens in European data centres too, not just storage. Both are aimed at Enterprise and Edu, are configured per project and go through sales rather than a toggle in your account. If you need it, you have to negotiate for it.

And the gap: Sites, the website builder OpenAI markets as part of Work, is unavailable in the EEA, Switzerland and the UK, and supports neither data nor inference residency. What that means for Europe is covered in ChatGPT Sites.

The wider point is the same as with any cloud service: an agent allowed to read inbox, storage and chat sees more than any single tool before it. Which content to entrust to someone else’s infrastructure at all is sorted out in ChatGPT and privacy.

Governance: who may do what – and who answers for the agent

OpenAI ships more tooling for workspaces than the headlines suggest: restrictions on plugins, browser and network access, an automatic review layer in which a strong model checks risky actions, a compliance interface for IT, and spend limits.

None of that replaces a rule of your own, and the most important one is uncomfortable: an agent that reads external content and is allowed to act is attackable through exactly that content. A doctored email, a manipulated document, a page with hidden instructions – the attack path is called prompt injection, it is architectural, and no vendor has solved it completely; the mechanics and the countermeasures that actually work are in prompt injection: securing AI agents properly.

Three guardrails cover most of the practice:

  1. Read broadly, write narrowly. Generous access to sources, minimal and explicit rights to change and send.
  2. Human sign-off for anything outbound. Emails, bookings, commits, publications. No exceptions, however tedious.
  3. Log it and spot-check it. What an agent did has to be reconstructable – otherwise the liability question has no answer.

That adoption fails on people rather than technology is the rule, not the exception – see AI acceptance in teams.

Work, Cowork, Copilot: where OpenAI sits in the field

Work is the answer to two products that arrived earlier. The three differ less in intelligence than in where they do their work:

Approach Runs Strength Available since
ChatGPT Work in the cloud, app-centric Breadth of connectors and output formats 9 July 2026
Claude Cowork on the desktop, on the local file system Long documents, work in real folders January 2026 (Mac), February (Windows)
Copilot Cowork inside the Microsoft 365 tenant Existing Microsoft stack and permissions generally available since 16 June 2026

Crowning a winner would be dishonest, because the decision rarely turns on the model. It turns on where the data lives that the agent is meant to work with: in the Microsoft tenant, on local drives, or spread across a dozen cloud services. The third pattern is the case Work was built for.

Who Work pays off for today

  • Teams with many systems and recurring reports. This is exactly where connectors plus scheduled tasks earn their keep.
  • Solo operators with a research and documentation load. An agent that turns scattered sources into a presentable first draft saves real hours.
  • Organisations already standardised on ChatGPT. The mode costs no additional contract, only allowance.

Anyone whose access permissions are not under control, who is unwilling to define approval steps for outbound actions, or who strictly needs European inference residency without negotiating an Enterprise agreement, should postpone adoption rather than force it.

The bottom line

ChatGPT Work is the most coherent version yet of what OpenAI has been signalling for two years: the assistant stops answering and starts delivering. The three model tiers make costs more predictable, voice finally removes the stop-start rhythm from the interaction – and the connectors are the part that decides whether any of it is useful.

The sober half: consumption is hard to predict, the voice API is still missing, and the European fine print – contract, residency, the blocked Sites feature – has to be actively clarified rather than assumed. If you start today, start with one recurring job, narrow write permissions and a person who reads the output. Exactly as the practical guide to AI adoption recommends for every stage before this one.

FAQ

Frequently asked questions

Do I need a new subscription for ChatGPT Work?

No. Work is not a separate tier but a mode inside the existing plans – prices have not changed. It is available on the paid plans except Free and Go: Pro, Pro Lite, Enterprise and Edu got it first, Plus and Business a few days later. You still pay, though: agent runs draw on your usage allowance, and considerably faster than ordinary chat.

Is ChatGPT Work available in the EU and the UK?

Yes – Work itself is part of the regular rollout and usable in supported regions including the EU. The exception is the Sites sub-feature, the built-in website builder: it is not available in the EEA, Switzerland or the United Kingdom at launch, with no announced timeline. For lawful use involving personal data you also need a data processing agreement – available for Business, Enterprise and the API, not for a private account.

What is the difference between ChatGPT Work and Codex?

Codex is the coding agent: repositories, pull requests, tests. Work is broader and workplace-oriented – it pulls context from connected apps and files and produces finished artefacts: documents, spreadsheets, presentations, reports, web apps. Both share the same usage allowance and sit side by side in the desktop app; the new voice mode works in either.

Which of the three GPT-5.6 models should I use?

Terra is the sensible default for everyday work. Luna pays off when speed and volume matter and the task stays simple – classifying, summarising, bulk processing. Sol is the flagship tier for heavy reasoning, large codebases and long autonomous runs; via the API it costs twice as much as Terra and five times as much as Luna on input. Rule of thumb: build with Terra, then move up or down only where it measurably helps.

Can I responsibly give an agent access to my inbox and drive?

Only with rules. An agent that reads external content and is also allowed to act is fundamentally attackable through exactly that content – this is not one vendor's weakness but the architecture. In practice: read permissions generous, write and send permissions narrow; anything going outward – emails, bookings, commits, publications – needs human sign-off. OpenAI supplies admin controls, an automatic review layer for risky actions and a compliance interface; the roles and approval logic remain yours to design.