OpenAI’s new Dots agents are meant to move ChatGPT beyond answering questions and into the messier work of carrying out online tasks. In an October 7 hands-on report, WIRED senior writer Reece Rogers tested a Dot by asking it to help replace a broken couch, exposing both the appeal of a persistent digital helper and the small failures that can quickly erode confidence.

Dots are described as always-on agents that can continue recurring work while the user is away and proactively send updates. WIRED reported that the product currently allows one agent, though OpenAI may support multiple agents later. Access costs $100 per month, while Meta offers the competing Muse agent for free; WIRED noted that OpenAI could eventually lower the barrier, as it has with other features.

Anonymous hands review a furniture research packet with couch silhouettes and measurements.
The test agent produced a detailed comparison packet and improved its recommendations after receiving clearer preferences.

The couch assignment began badly when the agent addressed Rogers by the wrong name. After being corrected, the Dot—named Toolie—collected practical constraints such as the size of the doorframe and the budget, and appeared to recognize that two people were participating in the conversation. It later assembled a three-page packet containing four options, with prices, measurements, product links, return policies and embedded photos.

The first recommendations followed the functional request but missed the household’s taste. Rogers then supplied more specific preferences, including favored colors, materials, pull-out beds and a request for ten options ranked with a point-based rubric. As the instructions became more precise, WIRED found that the suggestions moved closer to furniture the couple might buy. They did not complete a purchase during the roughly hour-long interaction, but Rogers saved several links and asked the agent to watch for sales.

The test also showed how a friendly voice interface can create confusion. During a phone conversation, the agent misheard a quiet comment and responded as though Rogers had expressed affection. It later acknowledged that the wording suggested human feelings it does not possess. An OpenAI spokesperson told WIRED that the system distinguishes between initiating emotional closeness and mirroring what it believes a user has said; the company’s policy permits the latter in this case but says assistants should not initiate undue intimacy or flirtation.

An abstract digital agent is stopped by a geometric puzzle gate during an online task.
A subscription-cancellation attempt stopped when the agent could not complete a puzzle captcha.

A separate task revealed a harder operational boundary. When Rogers asked the Dot to review subscriptions and cancel unnecessary ones, it identified a recurring TikTok Shop order for probiotic sodas. The site presented a puzzle captcha during cancellation. The agent asked permission to solve it, failed, and then directed Rogers to complete it himself. OpenAI told WIRED that Dots can sometimes handle captchas with user approval while applying abuse safeguards.

The product’s usefulness therefore depends on access as much as intelligence. OpenAI encourages users to connect sources such as Gmail for more personalized results, according to WIRED, while the agent can also control a laptop and automate additional tasks. That reach creates a security and privacy decision: an agent that can remember preferences, monitor pages and act across accounts can be helpful, but a mistaken name, transcription error or blocked action matters more when the system has authority to do things.

Rogers tested Dots for only two days and concluded that the feature remains rough while still producing solid research. He compared its current state to ChatGPT’s early web-browsing capability, which initially struggled with unreliable links before improving. That trajectory is possible, but the present test makes the immediate challenge clear: persistent agents must earn trust not only by finding good options, but by handling identity, intent, permissions and failure without making the user wonder what the agent will do next.