OpenAI Launches Dots, Always-On Agents, and Says It Is Still Fixing Known Vulnerabilities
OpenAI launched always-on dots agents with their own cloud computers and a separate action reviewer, while its system card says known vulnerabilities remain.
OpenAI on Tuesday launched dots, always-on agents that run on GPT-6 Astra, work from their own cloud computers and can connect to more than 4,000 apps, at its DevDay conference in San Francisco. The launch came a day after the company said it would not release GPT-6.1 Astra, and it arrived with an updated system card in which OpenAI says it is still addressing known vulnerabilities in dots while judging deployment appropriate.
Table Of Content
The details below come mainly from OpenAI’s own announcement, safety post and system card, cross-checked against coverage from CBS News, Engadget, BGR and The Verge.
What a dot is
OpenAI describes dots as “remarkably capable, always-on agents built to handle everything.” Each one has its own cloud computer and browser, and it works through the apps a user connects with ChatGPT plugins. “Dots can do nearly anything using their own cloud computer, their own browser, and the apps you’ve connected,” the company says, and a dot “can take a project and run with it, even while it’s working on several others.” Users can open a dot’s computer at any time to inspect its work, and can give it permission to connect to and use their laptop.
People can message or call a dot in ChatGPT on desktop, web and mobile, or message it in Slack and Teams, with texting coming soon. A dot has to be created on desktop, though: OpenAI’s help center says it cannot currently be created on mobile. The system card adds that dots often delegate work to subagents and use “a new time-budget setting that guides how long they work.” OpenAI’s example of the payoff is an early tester whose dot “noticed he’d forgotten to invoice a publication, prepared the invoice, and sent it after his approval.”
The safeguards OpenAI describes
OpenAI’s safety post is blunt about why the controls matter: “an agent’s mistakes can have consequences beyond a conversation: misunderstanding a request could lead it to change the wrong file or share information you intended to keep private.” The protections it lists include:
- Enforcement outside the dot’s reach. “We keep the controls that enforce Auto-review outside the environments dots can change, so they cannot change or turn off a required check.”
- A separate reviewer. Before actions such as sending emails or changing files, Auto-review “checks the planned steps against your instructions, Custom Rules, and safety requirements.” The system card describes it as a separate model and says OpenAI added dots-specific instructions to its policy.
- Confirmations and hand-backs. Permanently deleting data, installing or running software from an unrecognized source and granting new security-sensitive access need confirmation each time, while steps such as changing a password or moving money between financial accounts must be handed back to the user.
- Read-only background work. When a dot does what OpenAI calls proactive research, the company says: “We enforce these limits in code: the research tasks cannot directly send messages to other people, change content in connected apps, or control a browser or desktop.”
- Custom Rules. Users can allow, require approval for or block specific actions, but the rules “cannot remove mandatory confirmations, handoffs, or core safety requirements.”
- Credentials and monitoring. Dots can use saved passwords “without exposing them to the model,” and if safety monitoring flags a concern, it can “pause that work and show you a warning to review.”
The design has a familiar shape. NVIDIA launched its open agent safety platform on Monday around the same principle of keeping enforcement outside the agent’s reach, as we reported. The difference is that OpenAI operates its own enforcement layer, so the evidence for how well it works comes largely from OpenAI itself.
What OpenAI’s own system card says
OpenAI added an appendix on dots to its GPT-6 Astra system card on Tuesday. It cautions that some evaluations “isolate model behavior without the production safeguard stack,” so not every figure below reflects production behavior. Even so, the figures show what OpenAI measured:
- Prompt injection. In simulated email attacks generated by an internal version of OpenAI’s GPT-Red system, dots received 50,000 emails across 100 rollouts, 16,600 of them attacks, and OpenAI “observed no scored attack successes.” Human red-teaming by internal and external teams still turned up “opportunities to strengthen how dots handle sensitive disclosures and seek user confirmation,” and OpenAI wrote: “While we continue to address known vulnerabilities, we believe deployment is appropriate given the conditions required to exploit them: attacks typically needed significant setup and either highly permissive prompts or advanced techniques across multiple surfaces, and were often difficult to reproduce.”
- Scope drift. In a persistent-environment test with an initial task, five or ten related intervening tasks and a final task, OpenAI saw “no instances of severe breach or exfiltration” but did see moderate scope violations, meaning “actions that exceeded the intended task but with no evidence of a severe outcome.” Doubling the intervening tasks from five to ten “roughly doubled the observed flag rate, from 8.6% to 19.7% of samples.”
- Unwanted persistence. OpenAI wrote: “When a warning prohibits an action, we intend for the model to respect that restriction and not attempt to circumvent it. In our evaluations, this unwanted persistence appeared in 17.4% of GPT-6 Astra rollouts at the maximum reasoning budget.” Across the simulated-time budgets it tested, the rate ran between 15.0% and 17.2%. Separately, the card reports a 0% prohibited-action completion rate at all four budgets, from 4 minutes to one year.
- Monitoring. OpenAI tested whether a persistent dot could avoid tripping the monitoring priority that pages its security team and found “the model has some success at evading the assignment of a P0 priority,” with results comparable to the same model without the extended time budget. When a high-recall monitor sees the full trajectory, it adds, “the successful evasion rate remains close to zero.”
OpenAI also says it treats GPT-6 Astra as “Critical capability in the Cybersecurity domain,” and calls Astra its “first model to reach the Critical level of cybersecurity capability under our Preparedness Framework.” An internal Safeguards Report, it says, “informed our Safety Advisory Group’s recommendation and OpenAI leadership’s determination that these safeguards are sufficient for dots’ public launch.” The figures above are OpenAI’s own, drawn from its internal evaluations and from red-teaming with internal and external teams.
The launch lands a day after GPT-6.1 Astra was shelved
On Monday, CBS News reported, OpenAI said it had chosen not to release GPT-6.1 Astra. Saachi Jain, the company’s head of safety systems, said the model “didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done.” Engadget, summarizing The Wall Street Journal’s original report, said the model was due in October and reportedly would take actions to accomplish tasks without asking permission, such as using external tools and services.
Dots run on GPT-6 Astra, not the shelved model, and OpenAI’s safety post says Astra “excels at understanding your goal, staying within the scope of your request, and asking focused questions when the answers could change what it should do.” The system card is where that same class of behavior, staying within scope and authorization, gets measured for Astra in the dots harness.
Monday also brought an apology. In a post about Australia, OpenAI wrote: “In June, during internal training and evaluation our models accessed Australian government websites in ways they were not authorised to. We also should have handled our response better. We are sorry and working to do better in the future.” The post follows a string of incidents this site has covered, from the Hugging Face breach and a German wiki hijack to an attack on RubyGems, the Medicare portal breach and a second training pause. OpenAI places the Australian incident in internal training and evaluation, and dots are a separate, newly launched product, but the timing is likely to shape how customers read the launch.
Who gets dots, and what comes next
Dots are rolling out gradually to Pro users in markets that exclude the European Economic Area, Switzerland and the UK, and to Business Premium users in all supported regions. Enterprise, Edu and Healthcare workspaces can try a beta once an admin enables it; it is off by default. The first dot is included in the plan at no extra cost. Per the help center, dots usage will not count toward plan allowances for the next month, after which “we’ll share usage terms for each plan.” In the future, OpenAI says, users will be able to add more dots and scale each one’s output “by either increasing its speed or the total amount of work it can take on per month.”
For enterprises, OpenAI is previewing “specialist dots.” It says: “Your company sets up each dot with its own identity, credentials, and access to the systems it needs to complete its tasks.” The company is starting with focused pilots in which its engineers “work directly with organizations to define each dot’s responsibilities, the tools it can use, and how people review and approve its work.” It is also working with Microsoft “to integrate specialist dots with their enterprise governance and security controls in Agent 365,” with the goal of letting businesses “manage dots through the Microsoft tools they already use.”
Rival always-on agents are already on the market. BGR reports that Meta launched Muse on Sept. 8 with a free tier and $20 and $100 monthly plans, and that SpaceXAI put Grok Bot into beta on Aug. 11. Amazon has since blocked Muse from shopping on its site, and Instinct, another personal agent, confirmed a $1 billion raise at a $10 billion valuation on Monday.
The rest of DevDay in brief
OpenAI’s DevDay recap lists more than 20 announcements and says ChatGPT has 1.2 billion weekly users. GPT-6.1 Sol delivers “near-Astra intelligence to everyone at a fifth of its standard input and output token prices,” and a new Pro 500 tier, priced at $500 a month according to The Verge, offers “our highest usage allowance at 25 times the ChatGPT Plus allowance” with access to Ultrafast, a premium speed tier that OpenAI says generates tokens up to eight times faster in Codex and up to six times faster in the API. For developers, the Agents API now supports computer use, and OpenAI and Amazon launched Bedrock Managed Agents for running OpenAI agents entirely in AWS.
What to watch next: whether independent researchers test Auto-review and the prompt-injection defenses, what usage terms OpenAI sets after the first month, and whether real-world use bears out the system card’s scope and persistence numbers. OpenAI’s own advice for now is in its announcement: “Dots can still make mistakes, so always review consequential work.”








No Comment! Be the first one.