OpenAI is promoting agentic versions of its models—built on Codex and packaged as ChatGPT Work—to perform multistep, data-heavy office tasks. The product is available on the company’s lowest subscription tier for $20 a month and is intended to let language models access email, calendars and a range of SaaS tools to complete workflows autonomously.
Adoption gap and commercial stakes
An OpenAI-backed study found wide adoption differences in June: 98% of OpenAI employees were using Codex, while just 17% of organizational subscribers and under 1% of individual subscribers used the agentic coding tool. The company’s combined desktop and mobile app has about 20 million users, compared with more than a billion people who prompt ChatGPT online.
OpenAI engineers say longer-running agents consume more tokens and therefore increase per-user revenue, making wider uptake in non-engineering professions commercially important. Christian Catalini of a16z warned that if labs cannot secure complementary assets, value could accrue elsewhere.
Design trade-offs and technical limits
Engineers describe a tension between building a simple harness—software that decides what the model sees and which tools it can use—and relying on stronger general models. Gershenson, an engineering lead on the harness, said the team focuses on exposing only the information a model needs, arguing that future models may reduce the need for complex harnesses.
OpenAI staff acknowledge practical limits. Ambrosino, lead engineer on the desktop app, said agents must handle “the messy world” of legacy websites and workflows. The product currently has usability issues: permission flows for cloud drives can be confusing, some settings are available only on the web app, and certain actions (for example, creating events versus new calendars) are inconsistent.
Thibault Sottiaux, who leads core product work including Work, argued the conversational interface helps users, while critics such as Wharton professor Ethan Mollick suggested competing offerings like Claude favor iterative comparisons that ask for user input. OpenAI engineers also note that early adopters inside the company are informing product design and benchmarks drawn from a 44-occupation test suite called GDPval.
Engineers say expanding agents beyond developers will require better discoverability and clearer permission models. They emphasize incremental improvements to the harness while preparing for stronger models to handle more context directly.
Original source: TechCrunch AI