Agent framework
The agent framework, or harness, runs the loop: it sends context to the model, runs the tool call the model returns, feeds the result back, and repeats until the task is done. You either build your own with a framework, such as the Claude Agent SDK, the OpenAI Agents SDK, or the Vercel AI SDK, or use a finished agent that already has one, such as Claude Code, Codex, or Replit. A harness’s key pieces include:- Model: decides the next action from the system prompt, the conversation, tool results, and screenshots. It doesn’t drive the browser itself: it returns a tool call and the harness carries it out.
- System prompt: the instructions sent with every request to the model: its role, its rules, and the format of its answers.
- Tools: functions the model can ask the harness to call, like “navigate to this URL”, “click here”, or “run this code”.
- Skills: instructions and reference files the agent loads only when a task needs them, unlike the system prompt, which is sent every turn.
- Browser automation framework: turns tool calls into actions on the page, through the DOM with selectors (Playwright, Puppeteer) or through the screen with clicks at coordinates (computer use). Most agents combine the two: DOM actions for most steps, and computer use where the DOM doesn’t cooperate.
Browser infrastructure
Browser infrastructure is where the agent’s decisions actually get carried out. The agent framework decides to open a page, fill a form, or click a button; browser infrastructure runs the browser that does it, and gives the agent what it needs to reach and interact with real websites. Its key pieces include:- Browser: the Chromium instance the agent drives, isolated from every other session.
- Stealth and proxies: anti-detection, CAPTCHA handling, and an IP address that fits the site, so the browser isn’t blocked.
- Credentials and state: saved logins, credentials, and payment methods the browser can use without exposing them to the model.
- Observability: live view, replays, and telemetry, so you can see what happened when something goes wrong.
- Scale: running many browsers concurrently, and absorbing bursts when an agent creates a large number of browsers at once.