OpenAI’s new Dot agent appears most capable when it can work inside tools and files a user controls, while ordinary consumer websites still present stubborn obstacles. In a hands-on test published by The Verge, the agent struggled with security checks, payment steps, and several seemingly simple errands, but performed more convincingly when it was asked to edit media and rebuild a personal website.
OpenAI introduced Dots earlier in the week as customizable, blob-like agents, with one available initially and more planned. According to The Verge, a Dot operates through a virtual machine, can use applications including Blender and GIMP, and can connect to a user’s own computer through the ChatGPT desktop app. It also supports voice calls. Access is initially limited to OpenAI’s highest-priced account tiers, including the $100-per-month Pro plan.

The Verge’s test exposed the gap between the agent’s broad promise and the friction of the public web. Asked to arrange internet service, Dot found a promotional email offering a $100 incentive. It could not complete a press-and-hold verification challenge, however, and required the reviewer to take over. The agent later reached checkout, but the process called for bank information that the reviewer chose not to provide.
Other errands produced similar resistance. Dot was unable to secure a free tour of a coworking space and instead suggested a $35 day pass, while a competing agent called Instinct found a free trial, the report said. Dot also became stuck in repeated security checks while attempting tasks involving Ikea and a restaurant. The reviewer said these jobs required more manual intervention than comparable tests of Muse and Instinct.
That pattern matters because the obstacles were not necessarily failures of language understanding. They arose at the points where websites deliberately distinguish people from automation, require sensitive details, or depend on unpredictable interfaces. The test therefore suggests that an agent can be competent at planning and still be unreliable as an end-to-end consumer concierge when outside services impose their own gates.

Dot looked substantially stronger in a controlled working environment. The Verge reported that it redesigned the reviewer’s personal website, accepted iterative direction through both voice and text, and deployed the resulting changes. It also combined video files from the reviewer’s computer into a clip prepared for social media. The process demanded numerous permissions, but the requested work was completed.
Those successes frame Dot less as a universal web shopper than as a general-purpose work agent. The Verge characterized the experience as closer to Codex for nontechnical users, with an emphasis on business and production tasks. That interpretation fits the contrast in the test: websites owned by third parties introduced blockers, while the reviewer’s own files, software, and deployment workflow gave the agent a clearer path to act.
The value proposition remains uncertain. The reviewer did not find the tool worth $100 a month for their own work, though they suggested it could be more useful to someone launching a business. Dot eventually helped order a teriyaki meal after receiving assistance with account access, but the broader test indicates that dependable autonomy still depends heavily on the environment. For now, OpenAI’s agent looks most persuasive where users can grant explicit access and keep the workflow within systems they control.

Comments
Loading comments…