HomePortfolioBlogWorkContactGames

OpenAI DevDay 2026: AI Agents That Use Computers, Explained

Published on: 8th October, 2026

Hey there!

OpenAI held its DevDay 2026 in San Francisco on the 29th of September, and the headline was pretty clear: AI agents that use computers. Not just chatbots that answer questions, but agents that can open a browser, click buttons, fill in forms and keep working on a task after you close the tab.

I build frontends for a living and I run a small AI chat agent on this very site, so I wanted to break down what was announced, in plain English, and what I think it means for those of us who build things for the web.

First, what is an "AI agent"?

A chatbot takes your message and replies with text. An agent takes a goal and works towards it in steps. At each step it can decide to use a tool (look something up, call an API, save a record), look at the result, and decide what to do next. That back-and-forth is usually called the agent loop.

The "Ask Amin" chat widget on this site is a tiny version of this. When you ask it about one of my projects, the model can call a get_project_details tool that reads from a knowledge base, and if you are a recruiter it can call a tool that saves your details to a database. I cap that loop at five rounds, so it can never go off on an endless adventure. Every tool is something I wrote by hand, with a clear name and a clear job.

What OpenAI announced takes that same idea and gives the agent a much, much bigger tool: an actual computer.

The three DevDay announcements worth knowing

1. Dots: always-on agents. Dots are agents that run on their own cloud computer, can connect to more than 4,000 apps through plugins, and keep working on a task while you get on with your day. OpenAI's examples include investigating a bug, turning a design into a working app and preparing invoices. They are rolling out to the higher paid ChatGPT plans first, and you can check what your dot has been doing on its computer at any time.

2. Computer use in the Agents API. This is the developer side. The computer use tool lets an agent you build drive a browser inside a sandbox that OpenAI hosts for you. A sandbox is just an isolated environment, so whatever the agent does in there cannot touch your own machine. The Agents API also brings multi-agent workflows, tool search and something called context compaction, which is a way of squashing a long conversation down so the agent does not run out of memory halfway through a big task.

3. GPT-6.1 Sol: a cheaper workhorse model. GPT-6.1 Sol is pitched as getting close to OpenAI's top model on coding and computer use, at around one-fifth of the cost. API pricing is $2 per million input tokens and $10 per million output tokens, with "cached" input (text the model has already seen recently) at just $0.10 per million. Agents re-read a lot of the same context on every step, so that cached price matters more than it looks.

How does an AI agent actually use a computer?

This is the bit I found most interesting, because it is so simple. According to OpenAI's docs, it is a loop:

  • The agent takes a screenshot of the browser.
  • It looks at the screenshot and decides what to do next.
  • It clicks, types or scrolls.
  • It takes another screenshot, and repeats until it is done.

That's it. It is basically how you or I use a website, just a lot more literal. There is no special API or integration needed on the website's side. The agent sees the same page your users see.

Compare that to my chat agent. With hand-written tools, I decide exactly what the model is allowed to do. With computer use, the "tool" is the whole screen, and the agent figures out the rest. That is far more powerful, and also far harder to fence in.

What this means for frontend developers

Here is my take as someone who has spent more than ten years building React and TypeScript interfaces: your website is about to get a new kind of visitor.

If agents navigate by looking at the screen, then the things that make a page easy for a person to use probably make it easier for an agent too. A button that says "Continue to payment" is easier to reason about than an unlabelled icon. A form with clear labels and helpful error messages is easier to fill in correctly. A layout that jumps around while things load is confusing for everyone.

None of this is new advice. It is the same stuff accessibility and good UX have been asking for all along. The difference is that the payoff just got bigger, because some of your users may soon be agents working on behalf of a person.

I'd also expect more testing to go this way. A lot of my career has involved writing end-to-end tests with tools like Cypress, where you script every click by hand. OpenAI's own docs list testing a website as one of the use cases for computer use. I don't think it replaces proper test suites any time soon, but describing a user journey in plain English and letting an agent try it is a pretty appealing extra check.

The security part (please don't skip this)

A while back I did a short contract at IBM fixing issues found in a security test of a web application, covering things like IDOR (where changing an ID in a URL lets you see someone else's data) and people tampering with payment requests. The lesson that stuck with me: assume anything the client sends can be wrong or malicious, and check it on the server.

Agents make that lesson more important, not less. An agent clicking around your app is still a client. It can misread a page, follow instructions hidden in some text it shouldn't trust, or just make a mistake. Your backend should not care whether a click came from a human or an agent; the same permission checks need to hold.

OpenAI seems aware of this. For Dots, apps connect with read-only permissions by default, sensitive actions get an automatic review against the user's instructions, users can write custom rules to allow or block actions, and some things, like changing a password, stay human-only. On the API side, the agent has to ask before it visits each new website, sign-ins go through their own approval step, and the docs warn that screenshots can contain private account data, so you should keep them out of your logs.

Having led the frontend for a multi-chain crypto wallet, where a single wrong confirmation can move real money, I'm a big fan of that "ask a human before anything irreversible" pattern. If you are building with agents, design those approval moments in from day one rather than bolting them on later.

Should you start building with AI agents now?

If you already have a small AI feature, like a support bot or a chat widget, I would start with the boring wins: a cheaper model, and making better use of cached input. That alone can make a project that was too expensive to run suddenly reasonable.

Full computer use is exciting, but I would keep it to low-risk jobs to begin with: research, filling in internal forms, or checking that a sign-up flow still works. Give it its own accounts, limited access, and a human approval step for anything that can't be undone.

Takeaway

DevDay 2026 moved AI from "answers your questions" to "does tasks on a computer for you". For most of us in tech, the practical points are simple:

  • Agents work in a loop: look, decide, act, repeat. Computer use just makes the screen itself the tool.
  • Clear, accessible interfaces help both people and agents.
  • Treat agents like any other client: check permissions on the server and ask a human before anything irreversible.
  • Cheaper models with cached input make small AI features much easier to justify.

I'll be keeping my own chat agent on its short, hand-written leash for now, but I'm definitely going to try pointing a computer-use agent at this site to see how well it gets around. Stay tuned!