Overview
A chatbot answers questions. An agent does work. It opens a browser, logs into a portal, reads a spreadsheet, fills in a form, and tells you when it's done.
As of this fall, both leading AI labs offer serious tools for building these agents. At DevDay on September 29, OpenAI added computer use to its Agents API, letting developers build agents that interact with software to complete tasks. Anthropic offers the Claude Agent SDK, which gives developers the same tools, agent loop, and context management that power Claude Code.
This guide is for two readers: the business owner deciding whether an agent could take work off their team's plate, and the developer deciding how to build one.
What an AI agent can do for a business
Good agent tasks are repetitive, rule-based, and spread across several systems. Examples we see regularly:
- Pulling data from portals with no API. Insurance eligibility checks, vendor portals, government sites. The agent logs in, finds the record, and copies what you need.
- Processing incoming documents. Reading invoices, intake forms, or applications and entering the details into your system.
- Testing your own website or app. Walking through checkout or signup flows after every update to catch breakage.
- Research and monitoring. Checking bid boards, competitor pricing, or regulatory updates on a schedule and summarizing changes.
- Back-office follow-up. Drafting replies, updating CRM records, and flagging items for a human.
Bad agent tasks involve judgment calls with legal or financial consequences and no review step. Agents should prepare those decisions, not make them.
How an agent works, in plain English
Every agent, whatever the vendor, runs the same loop:
- Get a goal. "Find all open bids for web development posted this week."
- Look. Read the screen, the page, or the data it was given.
- Decide the next step. Click, type, search, call a tool, or ask a human.
- Act. Take that step inside a controlled environment.
- Check the result and repeat until the goal is met or it needs help.
The model provides the reasoning. Everything around it — the tools it can use, the permissions it has, the logs it keeps, the points where a human approves — is the engineering that makes an agent safe enough to use.
Option 1: OpenAI Agents API
The Agents API is OpenAI's hosted way to run agents. OpenAI says it now brings Codex's multi-agent capabilities, tool search, tool calling, and context compaction into your application, while OpenAI runs the underlying infrastructure.
For computer use, the agent works in an OpenAI-hosted browser. Your application starts a session, follows its progress, and approves each new website the agent wants to visit. The docs also describe a sign-in flow where users enter credentials through your interface, kept out of the model's input.
OpenAI also announced Bedrock Managed Agents with Amazon, so agents built this way can run entirely inside AWS.
Best for: teams that want hosted infrastructure, browser automation without managing their own sandbox, or an AWS-native deployment.
Option 2: Claude Agent SDK
The Claude Agent SDK is a code library for Python and TypeScript. Its docs describe it as giving you the same tools, agent loop, and context management that power Claude Code, with built-in tools for reading files, running commands, and editing code.
Because it's built on Claude Code's harness, it includes things production agents need: subagents for splitting work, project memory, skills, fine-grained tool permissions, and support for MCP servers that connect to your databases and APIs.
The SDK runs where you run it, which gives you control over the environment. Anthropic offers a separate hosted product, Managed Agents, for long-running agents when you don't want to manage sandboxes yourself.
On the model side, Anthropic reports Claude Sonnet 5.5 at 80.1% on OSWorld 2.1, a computer-use benchmark, close to Opus 5.5's 81.8%.
Best for: teams that want agents running in their own environment, coding and document-heavy agents, and deep customization of tools and permissions.
Choosing between them
| Question | OpenAI Agents API | Claude Agent SDK |
|---|---|---|
| Who runs the infrastructure? | OpenAI (or AWS via Bedrock Managed Agents) | You do (or Anthropic via Managed Agents) |
| Browser and computer use | Hosted browser with origin approvals and managed sign-in | Supported through Claude's tools in your environment |
| Languages | Official SDKs in several languages | Python and TypeScript |
| Connecting your systems | Functions, MCP connections, plugins | Custom tools and MCP servers |
| Strongest fit | Hosted browser automation, AWS deployments | Self-hosted agents, coding and file-based work |
You don't have to pick one forever. Because both support MCP, the connectors you build for your business systems can be reused. Many of our projects use one platform for the agent and keep the tool layer portable.
Five safety rules for production agents
OpenAI's own documentation is refreshingly blunt about the risks. Approving a website does not enforce confirmation before individual actions, and website content should be treated as untrusted. Those warnings apply to every agent platform.
- Least privilege. Give the agent only the accounts, tools, and data the task needs.
- Hard stops for consequential actions. Payments, deletions, and outbound messages need a human click, enforced in code, not just in the prompt.
- Treat everything the agent reads as untrusted. A webpage or email can contain text trying to redirect the agent. Your design must assume it will.
- Log everything. Keep a record of what the agent saw and did, without storing passwords or sensitive screenshots where they don't belong.
- Measure before scaling. Run the agent against past cases with known answers. Expand only when accuracy meets your bar.
Frequently asked questions
Do I need an agent, or just automation? If the steps never change and every system has an API, traditional automation is cheaper and more predictable. Agents shine when steps vary or a system can only be used through its screens.
Can an agent log into our systems safely? Yes, with proper design: dedicated accounts, limited permissions, and credentials entered through a secure flow, never pasted into the chat.
How long does a first agent take? A focused, single-workflow agent can reach a working prototype quickly. Hardening it for daily use is where most of the effort goes.
Start with one workflow
The best first agent takes a task your team does every week, hates doing, and can check easily. DDN builds agents on both platforms as part of our AI automation and custom software work, using the forward-deployed approach of building alongside your team on real cases.
Book a 15-minute discovery call and bring the workflow you'd most like to hand off.
Keep exploring
Related services
Need help applying this?
Design Develop Now builds websites, apps, and SEO-ready digital systems for businesses that need practical execution.
Start a project