Skip to main content
Introducing TSK-1Introducing TSK-1·Taskade's intelligence layer.
taskade
PricingHelpDashboard →Dashboard →
PricingLoginSign up for free →Sign up for free →
Dashboard →Dashboard →
Sign up →Sign up →
Loved by 1M+ users·Hosting 150,000+ apps·Deploying 500K+ AI agents·Running 1M+ automations·Backed by Y Combinator·Powered by TSK-1
TaskadeAppsTSK-1BuilderPricingFeaturesContact usIntegrationsMCP ServerPressAbout
ConnectProductivityVideosReviewsFAQ
LearnGenesisProjectsAI Agents
AutomationConnectorsAccount & BillingImport & ExportVideo TutorialsSearch Articles
DocsGetting StartedREST APIAction API
MCP ServersGuides & SDKModels
Community
FeaturedQuick AppsToolsDashboardsWebsites
WorkflowsProjectsFormsCreators
DownloadsMobile appsDesktop appsAndroidiOSMacWindows
ChromeFirefoxEdge
Compare
vs Cursorvs Boltvs Lovablevs V0vs Windsurf
vs Replitvs Emergentvs Devinvs Claude Codevs ChatGPTvs Claudevs Perplexityvs GitHub Copilotvs Figma AIvs Notionvs ClickUpvs Asanavs Mondayvs Trellovs Jiravs Linearvs Todoistvs Evernotevs Obsidianvs Airtablevs Basecampvs Mirovs Slackvs Bubblevs Retoolvs Webflowvs Framervs Softrvs Glidevs FlutterFlowvs Base44vs Adalovs Durablevs Gammavs Squarespacevs WordPressvs UI Bakeryvs Zapiervs Makevs n8nvs Jaspervs Copy.aivs Writervs Rytrvs Manusvs Crewvs Lindyvs Relevance AIvs Wrikevs Smartsheetvs Monday Magicvs Codavs TickTickvs Any.dovs Thingsvs OmniFocusvs MeisterTaskvs Teamworkvs Workfrontvs Bitrix24vs Process Streetvs Toggl Planvs Motionvs Momentumvs Habiticavs Zenkitvs Google Docsvs Google Keepvs Google Tasksvs Microsoft Teamsvs Dropbox Papervs Quipvs Roam Researchvs Logseqvs Memvs WorkFlowyvs Dynalistvs XMindvs Whimsicalvs Zoomvs Remember The Milkvs Wunderlist
Taskade AIVideo GuideAI App BuilderVibe CodingAgent BuilderDashboard Builder
CRM BuilderWebsite BuilderForm BuilderWorkflow AutomationWorkflow BuilderBusiness-in-a-BoxAI for MarketingAI for Developers
AI Agents
FeaturedProject ManagementOperations IntelligenceProductivityMarketing
TranslatorContentWorkflowResearchPersonalSalesSocial MediaTo-Do ListCRMTask AutomationCoachingCreativityTask ManagementBrandingFinanceLearning and DevelopmentBusinessCommunity ManagementMeetingsAnalyticsDigital AdvertisingContent CurationKnowledge ManagementProduct DevelopmentPublic RelationsProgrammingHuman ResourcesE-CommerceEducationLegalEmailSEODeveloperVideo ProductionDesignFlowchartDataPromptNonprofitAssistantsTeamsCustomer ServiceTrainingTravel PlanningUML DiagramER DiagramMath TutorLanguage LearningCode ReviewerLogo DesignerUI WireframeFitness CoachLead EnrichmentFounder OSSales DevelopmentBookkeepingRecruitingWebsite MonitoringField ServiceLicensingAll Categories
Automations
FeaturedAI Agent AutomationAI WorkflowsLogic AutomationsTrigger Automations
Agentic Process AutomationAction AutomationsAI Models in WorkflowsAgentic AutomationMulti-Agent AutomationBusiness-in-a-BoxOperations IntelligenceInvestor OperationsEducation & LearningHealthcare & ClinicsReal EstateStripeSalesHR & People OpsField Service & DispatchRenewals & LicensesE-commerceContentMarketingEmailCustomer SupportHubSpotProject ManagementAgentic WorkflowsAppointment SchedulingCalendarReportsSlackWebsiteFormTaskWeb ScrapingWeb SearchChatGPTText to ActionYoutubeLinkedInTwitterGitHubDiscordMicrosoft TeamsWebflowIndustry News & RSS FeedsGoogle WorkspaceManufacturing & OperationsAI Agent TeamsNotion AutomationsProposalBookkeeping & ExpensesClient OnboardingGoogle SheetsGoogle DriveGoogle CalendarGoogle FormsShopifyAsanaAirtableTrelloTodoistMailchimpClickUpGoogle DocsGmailGoogle TasksJiraLinearMicrosoft OutlookTelegramCalendlyTypeformSalesforceMonday.comTwilioWhatsApp BusinessWordPressZoomHTTP and WebhooksMCP AutomationFacebook PagesRedditApolloAll Categories
Wiki
Taskade GenesisAI AgentsAutomation
ProjectsLiving DNAAutonomous Workspaces, Agents & AppsQuantum AI & Taskade Genesis QuantumFoundations: The Theory Under Modern AIPlatformAI InfrastructureIntegrationsProductivityMethodsProject ManagementAgileScrumAI ConceptsCommunityTerminologyFeatures
Templates
FeaturedChatGPTOperations IntelligenceTablePersonal
Project ManagementSalesFlowchartTask ManagementEngineeringEducationDesignTo-Do ListMarketingMind MapGantt ChartOrganizationalPlanningMeetingsTeam ManagementStrategyGamingProductionProduct ManagementStartupRemote WorkY CombinatorRoadmapCustomer ServiceLegalEmailBudgetsContentConsultingE-CommerceStandard Operating Procedure (SOP)Human ResourcesProgrammingMaintenanceCoachingSocial MediaHow-TosResearchMusicTrip PlanningCRMClient OnboardingEmployee OnboardingSOPBug TrackerRecruitment TrackerFormSales PipelineContent CalendarMarketing PlanProduct RoadmapBusiness PlanSWOT Analysis30-60-90 Day PlanInterviewNotion AlternativeKPIStrategic PlanMeeting AgendaInvoiceRisk RegisterIT Asset ManagementKanban BoardChange ManagementCommunication PlanRFPScope of WorkStatement of WorkHelpdeskKnowledge BaseCreative BriefGoal SettingExecutive SummaryGap AnalysisBooking SystemEvent ManagementPortfolio TrackerCustomer Onboarding PortalsClient PortalAgency OperationsFinance TrackingAll Categories
Generators
AI SoftwareNo-Code AI AppAI AppAI WebsiteAI Dashboard
AI FinanceAI Operations IntelligenceAI FormAI AgentAI Client Portal BuilderAI WorkspaceAI ProductivityAI To-Do ListAI WorkflowsAI EducationAI Mind MapsAI FlowchartAI Scrum Project ManagementAI Agile Project ManagementAI MarketingAI Project ManagementAI Social Media ManagementAI BloggingAI Agency WorkflowsAI ContentAI Software DevelopmentAI MeetingAI PersonasAI OutlineAI SalesAI ProgrammingAI DesignAI FreelancingAI ResumeAI Human ResourceAI SOPAI E-CommerceAI EmailAI Public RelationsAI InfluencersAI Content CreatorsAI Customer ServiceAI BusinessAI PromptsAI Tool BuilderAI SEOAI Gantt ChartAI CalendarsAI BoardAI TableAI ResearchAI LegalAI ProposalAI Video ProductionAI Health and WellnessAI WritingAI PublishingAI NonprofitAI DataAI Event PlanningAI Game DevelopmentAI Project Management AgentAI Productivity AgentAI Marketing AgentAI Personal AgentAI Business and Work AgentAI Education and Learning AgentAI Task Management AgentAI Customer Relations AgentAI Programming AgentAI SchemaAI Business PlanAI Pitch DeckAI InvoiceAI Lesson PlanAI Social Media CalendarAI API DocumentationAI Database SchemaAI Marketing PlanAI Sales Pipeline GeneratorAI Course BuilderInternal ToolsBooking SystemReal Estate CRMInventory ManagementAI CRM BuilderAI TimesheetAI DispatchAI NewsletterAI Clinic OperationsAI Directory BuilderAll Categories
Converters
AI Featured ConvertersAI PDF ConvertersAI CSV ConvertersAI Markdown ConvertersAI Prompt to App Converters
AI Data to Dashboard ConvertersAI Workflow to App ConvertersAI Idea to App ConvertersAI Flowcharts ConvertersAI Mind Map ConvertersAI Text ConvertersAI Youtube ConvertersAI Knowledge ConvertersAI Spreadsheet ConvertersAI Email ConvertersAI Web Page ConvertersAI Video ConvertersAI Coding ConvertersAI Task ConvertersAI Kanban Board ConvertersAI Notes ConvertersAI Education ConvertersAI Language TranslatorsAI Business → Backend App ConvertersAI File → App ConvertersAI SOP → Workflow App ConvertersAI Portal → App ConvertersAI Form → App ConvertersAI Schedule → Booking App ConvertersAI Metrics → Dashboard ConvertersAI Game → Playable App ConvertersAI Catalog → Directory App ConvertersAI Creative → Studio App ConvertersAI Agent → Agent App ConvertersAI Audio ConvertersAI DOCX ConvertersAI EPUB ConvertersAI Image ConvertersAI Resume & Career ConvertersAI Presentation ConvertersAI PDF to Spreadsheet ConvertersAI PDF to Database ConvertersAI PDF to Quiz ConvertersAI Image to Notes ConvertersAI Audio to Notes ConvertersAI Email to Tasks ConvertersAI CSV to Dashboard ConvertersAI YouTube to Flashcards ConvertersURL to NotesVideo → SummaryAI Receipts to Expense Tracker ConvertersAI Docs to Knowledge Base ConvertersAI Form to Client Portal ConvertersSpreadsheet to CRMAll Categories
Prompts
Client PortalsDocument IntakeField Service
Operations IntelligenceCRMAdvertisingDesignInvoicing and PaymentsSocial MediaQuotes and EstimatesCalculatorsCodingBooking and SchedulingConsultingCustomer SupportComplianceTime TrackingRetail OperationsSaaS and SubscriptionsInsuranceGovernance and OversightNonprofitPatient IntakeBlog WritingBrandingPersonal FinanceHuman ResourcesPublic RelationsTeam CollaborationProduct ManagementSupportAgencyReal EstateMarketingResearchSalesCopywritingContentProject ManagementWebsite CreationStrategyE-commerceEngineeringSEOEducationEmail MarketingUX/UIProductivityInfluencer MarketingAnalyticsEntrepreneurshipLegalVibe CodingRecruitingAll Categories
Blog
Introducing Taskade TSK-1: The System Kernel Behind Every App (2026)AI Cost per Task (2026): What AI Work Really Costs to RunA Year of AI Agent Memory Experiments: Four Negative Results (2026)
What Is GPT? GPT vs LLM vs ChatGPT, and Why Models Come in Tiers (2026)Clients, Projects, Teams: Your Workspace, Now an App (September 2026)Agent Handoff Explained: The Narrow Channel Between AI Agents (2026)The Complete Connectome Explained: What a Finished Brain Map Tells Us About AI Agents (2026)RL Environments Explained: How AI Labs Now Train Models on Real Work (2026)Which AI App Builder Remembers Why? 10 Tools Compared for Month Six (2026)The TSK-1 Methodology: How We Benchmark AI Models by Building Real Apps (2026)Paste Your Website, Clone an App Kit, Launch It Live (September 2026)Open Source Temporal Parser: Parse ISO 8601, RFC 3339, and IXDTF in TypeScript (2026)Your App, Your Domain: Build CRMs, Client Portals, and Shops (September 2026)Best AI Model for Building Apps in 2026: One Prompt, Nine Models, Real Apps Side by SideConnect It, Start It, Watch It Run: Every Automation Run in View (September 2026)Chat-Native App Builders in 2026: What You Actually Own When the Chat EndsGenerate the Art. Preview the Agent. Put It on Your Domain (2026)Agentic Automation Explained: Agent vs AI Step (2026)The Scaffolding Tax: Why Less Prompt Beats More (2026)Run Your Whole Business in One App with Taskade Genesis (June 2026)
AIAutomationProductivityProject ManagementRemote WorkStartupsKnowledge ManagementCollaborative WorkUpdates
Changelog
Calendar RSVP Status & Workspace DNA Maps (Sep 25, 2026)Taskade Genesis Publish Hotfix (Sep 24, 2026)Calendar Attendees & Automation Setup Checks (Sep 24, 2026)
Taskade EVE Media Automations & Live Updates (Sep 22, 2026)Media Actions for Automations & Queued Messages (Sep 21, 2026)Large Project Import Stability Hotfix (Sep 21, 2026)Large Project Import Hotfix (Sep 21, 2026)
Wiki
Taskade GenesisAI AgentsAutomation
ProjectsLiving DNAAutonomous Workspaces, Agents & AppsQuantum AI & Taskade Genesis QuantumFoundations: The Theory Under Modern AIPlatformAI InfrastructureIntegrationsProductivityMethodsProject ManagementAgileScrumAI ConceptsCommunityTerminologyFeatures
Prompts
Client PortalsDocument IntakeField Service
Operations IntelligenceCRMAdvertisingDesignInvoicing and PaymentsSocial MediaQuotes and EstimatesCalculatorsCodingBooking and SchedulingConsultingCustomer SupportComplianceTime TrackingRetail OperationsSaaS and SubscriptionsInsuranceGovernance and OversightNonprofitPatient IntakeBlog WritingBrandingPersonal FinanceHuman ResourcesPublic RelationsTeam CollaborationProduct ManagementSupportAgencyReal EstateMarketingResearchSalesCopywritingContentProject ManagementWebsite CreationStrategyE-commerceEngineeringSEOEducationEmail MarketingUX/UIProductivityInfluencer MarketingAnalyticsEntrepreneurshipLegalVibe CodingRecruitingAll Categories
© 2026 Taskade
PrivacyTermsSecurity
Made withTaskade AIforBuilders
BlogAIBrowser Agents Explained: How…

Browser Agents Explained: How AI Agents Use the Web in 2026

What browser agents are, how they work, which ones shut down in 2026, and when to use one. The API-first, browser-last ladder, costs, risks, and agent identity.

Browser agents explained: the API-first, browser-last ladder for AI agents, from web search and page extraction to APIs and a real browser
November 17, 202628 min readStan ChangAI·#browser-agents#ai-agents#computer-use
On this page (19)
What Is a Browser Agent?How Does a Browser Agent Work?How an agent "sees" a pageA browser agent run, step by stepThe Four Kinds of Browser Agents in 2026What Happened to Browser Agents in 2025 and 2026?Why the first wave foldedWhen Should You Use a Browser Agent? The API-First, Browser-Last LadderHow Much Does a Browser Agent Cost?Are Browser Agents Safe? Prompt Injection and PermissionsHow Do Websites Know an Agent Is an Agent?How Reliable Are Browser Agents? What the Benchmarks SayHow to Build or Choose a Browser AgentWhat building with Stagehand looks likeWhere Taskade Fits: The Workspace Around the AgentKey Takeaways🔗 Related Reading📚 Sources💬 Frequently Asked Questions About Browser Agents

In 2025, every major AI lab shipped an agent that could use a web browser. By August 2026, OpenAI had retired two of them, Google had shut down a third, and Microsoft had announced the end of a fourth.

Browser agents did not die. They moved. The standalone "AI that browses for you" products were folded into bigger products, and the capability itself became infrastructure: cloud browsers, open-source frameworks, and signed identities that let a website know which agent is knocking. At the same time, builders learned an expensive lesson. Most of the web work an agent does should never touch a browser at all.

This guide explains what a browser agent is, how one works step by step, what happened to the famous ones, when to use one, and when an API, a search, or a simple page fetch will do the job better. 🧭

TL;DR: A browser agent is an AI that drives a real web browser: it reads the page, decides, and clicks. In 2026 the rule is API-first, browser-last: search, then fetch, then an API, and only then a browser. OpenAI retired Operator, ChatGPT agent, and Atlas by August 2026. See what AI agents can do without a browser →

What Is a Browser Agent?

A browser agent is an AI system that operates a real web browser to complete a task. A language model reads the page, picks the next action, and a runtime clicks, types, scrolls, or navigates. The loop repeats until the goal is met. Browser agents reach the parts of the web that have no API: portals, dashboards, forms, and checkouts.

The idea is simple, and the words around it are easy to mix up. This table separates the five terms people use interchangeably.

Term What it is Who decides the next step Example
Headless browser A real browser engine running without a visible window Code you wrote Chrome in headless mode
Browser automation script Fixed steps that drive a browser The script A Playwright test
Web scraper Code or a service that reads pages and extracts data The script or an extraction model A crawl API
Browser agent A model that chooses actions inside a browser A language model, step by step Claude in Chrome, a Stagehand agent
Computer-use agent A model that operates a whole desktop or phone A language model reading screenshots Anthropic computer use on a virtual machine

A browser agent is to a scraper what a driver is to a map. The scraper knows where the data sits. The agent can find its way when the road changes. That flexibility is the whole point, and also the whole cost: every decision is a model call.

For a shorter definition, see the browser agents glossary entry, and for the wider category, computer-use agents.

How Does a Browser Agent Work?

A browser agent runs a loop with four steps: observe the page, decide the next action with a language model, act in the browser, and check the result. It repeats until the task is done, a step budget runs out, or a guardrail stops it. Every step is a model call, which is why step count drives both speed and cost.

not done done risky step:payment, delete, send Goal from a personor another agent Observepage text, accessibility tree,or a screenshot Decidethe model picks one action Actclick, type, scroll, navigate Checkdid the page changeas expected? Return the result Ask a human
not done done risky step:payment, delete, send Goal from a personor another agent Observepage text, accessibility tree,or a screenshot Decidethe model picks one action Actclick, type, scroll, navigate Checkdid the page changeas expected? Return the result Ask a human

How an agent "sees" a page

The observe step is where browser agents differ most. There are three ways to show a page to a model, and each trades accuracy against cost.

Perception What the model receives Strength Weakness
Raw HTML or text The page's markup or visible text Cheap and fast Noisy, huge on modern sites, misses what is visible
Accessibility tree The structured list of buttons, links, and fields that screen readers use Compact, labels actions clearly Weak on canvas-heavy or unlabeled pages
Screenshot An image of the rendered page Sees exactly what a person sees Every step sends an image, so it is the slowest and costliest

Most production frameworks now blend the first two and fall back to screenshots. Browserbase credits the switch to the accessibility tree for a jump in Stagehand's reliability in early 2025. Pure screenshot agents, which click by coordinates, are the most general and the least precise. Browserbase founder Paul Klein IV put it plainly on the Latent Space podcast: coordinate clicking "proved to be less reliable than I would like," while anchoring each action to the actual page element means "it's more accurate."

A browser agent run, step by step

Find the cheapest SFO to Dubai round trip for May 8 to 15 open the flight search page accessibility tree of the page goal + page + history type SFO in the origin field type SFO updated page goal + page + history open the date picker, choose May 8 click, click results page extract the three cheapest flights as JSON flights with airline, price, stops result, plus a recording of the session repeats for each field, then submits You Agent loop Language model Browser
Find the cheapest SFO to Dubai round trip for May 8 to 15 open the flight search page accessibility tree of the page goal + page + history type SFO in the origin field type SFO updated page goal + page + history open the date picker, choose May 8 click, click results page extract the three cheapest flights as JSON flights with airline, price, stops result, plus a recording of the session repeats for each field, then submits You Agent loop Language model Browser

That flight search is not a toy. It is one of the harder public demos because date pickers, pop-ups, and dynamic results break scripted automation. In a sponsored 2026 tutorial, the YouTube creator Tech With Tim ran almost exactly this task with Browserbase's Stagehand framework in about 90 lines of code. The run took about four minutes, and the replay shows the agent drifting to the wrong page and then recovering.

The Four Kinds of Browser Agents in 2026

Browser agents come in four forms: assistants that live inside your own browser, agents that run a browser in the cloud for you, developer frameworks for building your own, and the cloud-browser infrastructure underneath. Most of the 2025 headlines were about the first two. Most of the 2026 growth is in the last two.

Kind Examples (September 2026) Runs where Built for
In-browser assistant Claude in Chrome, Perplexity Comet, Chrome auto browse Your own browser, logged in as you Individuals
Hosted cloud agent The cloud browser in ChatGPT Work, Browserbase Agents A browser in the provider's cloud Individuals and teams
Developer framework Stagehand, Browser Use, Playwright MCP Your code, local or cloud Developers building their own agents
Cloud-browser infrastructure Browserbase, Kernel, Anchor Browser, Steel, Hyperbrowser, Cloudflare Browser Run The provider's fleet of browsers Developers running many sessions
Runs in your browser Runs in a cloud browser for you Frameworks you build with Cloud-browser infrastructure Claude in Chrome Perplexity Comet Chrome auto browse ChatGPT Work cloud browser Browserbase Agents Stagehand Browser Use Playwright MCP Browserbase Kernel Anchor Browser Cloudflare Browser Run
Runs in your browser Runs in a cloud browser for you Frameworks you build with Cloud-browser infrastructure Claude in Chrome Perplexity Comet Chrome auto browse ChatGPT Work cloud browser Browserbase Agents Stagehand Browser Use Playwright MCP Browserbase Kernel Anchor Browser Cloudflare Browser Run

The split matters for risk. An assistant inside your browser acts with your cookies and your logins, so it can do anything you can do. A cloud browser starts clean, with only the credentials you hand it. That difference decides which tasks you can safely delegate.

The infrastructure layer has a history of its own. The history of Browserbase traces how one company went from a founder's memo in November 2023 to a $300 million valuation by building browsers for agents.

What Happened to Browser Agents in 2025 and 2026?

Between late 2024 and August 2026, most first-generation consumer browser agents were launched, then folded into larger products or shut down. OpenAI retired Operator, ChatGPT agent, and the Atlas browser. Google shut down Project Mariner. Microsoft announced the end of Copilot Mode in Edge. The capability survived inside bigger products and as developer infrastructure.

Product Company Launched What happened
Computer use Anthropic Oct 22, 2024 Became a standard model capability that many agents build on
Project Mariner Google Dec 11, 2024 Shut down May 4, 2026. Its technology moved into Gemini.
Operator OpenAI Jan 23, 2025 Folded into ChatGPT agent (Jul 17, 2025). The standalone site closed Aug 31, 2025.
ChatGPT agent OpenAI Jul 17, 2025 No longer available by July 2026. OpenAI points users to ChatGPT Work and its cloud browser.
Comet Perplexity Jul 9, 2025 Free worldwide from Oct 2, 2025. Android followed in November 2025 and iOS in March 2026.
Claude in Chrome Anthropic Aug 26, 2025 (pilot, as Claude for Chrome) Generally available on every paid Claude plan from Aug 26, 2026
ChatGPT Atlas OpenAI Oct 21, 2025 Deprecated. Stopped working Aug 9, 2026, with agentic browsing moved into ChatGPT and Codex.
Copilot Mode in Edge Microsoft 2025 Retired May 13, 2026. Its features moved into regular Edge, and agentic browsing became Browse with Copilot.
Chrome auto browse Google Jan 28, 2026 (U.S. AI Pro and Ultra, desktop) Reached U.S. Android subscribers on Aug 18, 2026
ChatGPT Work cloud browser OpenAI Jul 9, 2026 (paid plans except Free and Go) Website sign-in added for Plus and Pro on Aug 25, 2026
2024 2025 2026 Oct: Anthropic computer use Dec: Google Project Mariner Jan: OpenAI Operator Jul: ChatGPT agent + Perplexity Comet Aug: Claude for Chrome previewCloudflare signed agents Oct: ChatGPT Atlas, Gemini computer use May: Mariner shut down,Edge Copilot Mode retirement Jul: ChatGPT agent removed Aug: Atlas stops working,Claude in Chrome GA
2024 2025 2026 Oct: Anthropic computer use Dec: Google Project Mariner Jan: OpenAI Operator Jul: ChatGPT agent + Perplexity Comet Aug: Claude for Chrome previewCloudflare signed agents Oct: ChatGPT Atlas, Gemini computer use May: Mariner shut down,Edge Copilot Mode retirement Jul: ChatGPT agent removed Aug: Atlas stops working,Claude in Chrome GA

Why the first wave folded

Three forces explain the retreat, and none of them is "browser agents do not work."

  1. A standalone agent is a hard product to love. People do not want an AI that browses. They want the ticket booked, the form filed, the report finished. OpenAI's own explanation for retiring Atlas was that it was "moving browser-based agentic capabilities into ChatGPT and Codex", where the work already happens.
  2. Browsing is the slowest path to most answers. A model that can search, read, and call APIs gets most jobs done faster than one that clicks. As tool use improved, fewer tasks needed a browser at all.
  3. Trust and safety costs are real. An agent that acts with your logged-in session is exposed to prompt injection on every page it reads. Every lab that shipped one also shipped warnings.

What survived tells you where the value is: assistants embedded in the browsers people already use, cloud browsers inside larger work products, and developer infrastructure for companies that build the agent into their own software.

When Should You Use a Browser Agent? The API-First, Browser-Last Ladder

Use a browser agent only when nothing cheaper works. Try the rungs in order: search the web, fetch and extract a page, call an API or integration, and only then drive a real browser. Each rung is slower, more expensive, and less reliable than the one below it. Most agent web work never needs the top rung.

Paul Klein IV, whose company sells cloud browsers, described the same order for data collection on Latent Space: first a plain curl request, then a scraping API, and "if those two don't work, bring out the heavy hitter." When the company that sells the heavy hitter tells you to try something lighter first, it is worth listening.

yes no yes no yes no yes no, a desktop app What does the task need? Just an answerfrom public sources? Rung 1: web search Data on a public page? Rung 2: fetch and extract Does the service havean API or integration? Rung 3: API call or integration Web interface only:login, forms, clicks? Rung 4: browser agent Rung 5: computer-use agent
yes no yes no yes no yes no, a desktop app What does the task need? Just an answerfrom public sources? Rung 1: web search Data on a public page? Rung 2: fetch and extract Does the service havean API or integration? Rung 3: API call or integration Web interface only:login, forms, clicks? Rung 4: browser agent Rung 5: computer-use agent
Rung Method Typical speed Relative cost Main failure
1 Web search Seconds Lowest Stale or thin results
2 Fetch and extract a public page Seconds Low JavaScript-only content, bot walls
3 API or integration Milliseconds to seconds Low No API exists, or it misses fields
4 Browser agent Tens of seconds to minutes High: browser time plus a model call per step Wrong clicks, layout changes, CAPTCHAs, logins
5 Computer-use agent on a full desktop Minutes Highest Everything in rung 4, plus the operating system
  The cost of each rung, roughly

rung 1 search ▏▌ one query
rung 2 fetch + extract ▏▌▌ one page, one model pass
rung 3 API call ▏▌ one request
rung 4 browser agent ▏▌▌▌▌▌▌▌▌▌▌▌▌▌▌▌▌▌ browser time + a model call per step
rung 5 computer use ▏▌▌▌▌▌▌▌▌▌▌▌▌▌▌▌▌▌▌▌▌▌ a whole machine + a screenshot per step

Rungs 2 and 3 can swap places. When a service offers an official API for the same data, use it before scraping its pages: it is more stable, it survives redesigns, and it is the route the service wants you to use.

The ladder also explains a pattern in the market. Klein argues that most "computer use" is really browser use, and that you rarely need a whole operating system to click a button: "Do we really need to run an entire operating system just to control a browser? I don't think so." His estimate is that a browser-only fleet delivers "90% of the functionality… at 10% of the cost of running a full OS." Treat the numbers as a founder's claim. The direction matches what builders report.

How Much Does a Browser Agent Cost?

A browser agent costs browser time plus a model call on every step, plus retries. A 20-step task means about 20 model calls, often with page content or screenshots attached, and a browser session that stays open the whole time. Reading a page once is far cheaper. That is why step count, not model price alone, decides the bill.

Cost driver What moves it How to reduce it
Model calls Steps per task, and whether each step sends a screenshot Use the accessibility tree, keep steps atomic, cache actions that worked
Browser time How long each session stays open Close sessions promptly, avoid long waits
Retries Flaky pages, pop-ups, rate limits Handle known pop-ups in code, back off politely
Proxies and identity Sites that block data-center traffic Prefer signed identity and official APIs where they exist

Public price points give a sense of scale. Browserbase's Developer plan includes 100 browser hours for $20 a month, and its Fetch API costs about $1 per 1,000 pages, so reading a page is orders of magnitude cheaper than driving one. Caching helps too: Browserbase says server-side caching of Stagehand actions makes repeat workflows "up to 2x faster" with "~30% cost reduction."

Are Browser Agents Safe? Prompt Injection and Permissions

Browser agents are safe enough for many tasks with the right guardrails, but they carry one risk that ordinary automation does not: prompt injection. Any text on a page the agent reads can contain instructions, and a model can mistake them for yours. An agent logged in as you can then act on them.

The evidence comes from the vendors themselves. When Anthropic piloted Claude for Chrome in August 2025, its red-team tests (123 test cases across 29 attack scenarios, in autonomous mode) found that safety measures cut the prompt-injection attack success rate "from 23.6% to 11.2%", and to 0% on a set of browser-specific attacks. The same month, Brave showed that Perplexity's Comet "feeds a part of the webpage directly to its LLM without distinguishing between the user's instructions and untrusted content from the webpage." And the day after Atlas launched in October 2025, OpenAI's chief information security officer, Dane Stuckey, wrote that "prompt injection remains a frontier, unsolved security problem." By December, OpenAI's own guidance said prompt injection "is unlikely to ever be fully 'solved'."

The honest reading: mitigations work, and no vendor claims the risk is gone.

Guardrail What it prevents Example
Scoped accounts An agent with more access than the task needs A service login instead of your admin account
Domain allow-lists Wandering onto a malicious page Only the supplier portal and the shipping site
Confirmation gates Irreversible mistakes Ask before paying, deleting, or sending
Credential vaults Passwords reaching the model 1Password's agentic autofill fills logins without exposing them to the agent
Live view and recordings Silent wrong turns A person watches the first runs, and every session is replayable
A clean cloud browser Access to your personal logins Start each run without your cookies

Treat a browser agent like a capable new contractor: give it the smallest badge that does the job, watch its first shifts, and widen access as it earns trust.

How Do Websites Know an Agent Is an Agent?

Websites increasingly identify agents by cryptographic signature rather than by guessing. Web Bot Auth lets an agent platform sign each HTTP request, and a site or its CDN verifies the signature against a published key. Cloudflare launched signed agents on this basis in August 2025, and bot-detection vendors and payment networks adopted it.

For most of the web's history, automation hid. Scrapers rotated proxies, faked browser fingerprints, and paid CAPTCHA-solving services, and websites fought back with ever stronger bot detection. The arms race broke in 2025.

Date Event
Jul 1, 2025 Cloudflare blocks AI crawlers by default on new domains and launches pay per crawl
Aug 28, 2025 Cloudflare launches signed agents on Web Bot Auth. Its first cohort: ChatGPT agent, Goose from Block, Browserbase, and Anchor Browser.
Sep 2025 Cloudflare publishes the Content Signals Policy for robots.txt. Stytch adds Web Bot Auth support.
Oct 2025 Visa's Trusted Agent Protocol builds on the same signatures. The IETF forms a webbotauth working group.
Apr 2026 Browserbase: "What we previously framed as stealth is more accurately described as identity."
Jul 1, 2026 Cloudflare begins evolving pay per crawl into "Pay Per Use" and tests a use= field for Content Signals
Sep 15, 2026 New Cloudflare domains that show ads get separate defaults: search allowed, AI training disallowed, and agents blocked on pages with ads
alt [verified signed agent] [unsigned or invalid] request + Signature headers (RFC 9421) fetch the platform's public key key verify the signature allow, apply the site's agent policy page challenge, rate limit, or block Agent platform Website's CDN Published key directory Website
alt [verified signed agent] [unsigned or invalid] request + Signature headers (RFC 9421) fetch the platform's public key key verify the signature allow, apply the site's agent policy page challenge, rate limit, or block Agent platform Website's CDN Published key directory Website

Cloudflare's own taxonomy shows how seriously publishers now treat agents. Its Agent category covers "chat fetch bots (e.g., ChatGPT-User) and browser-use agents (e.g., Gemini or Claude driving Chrome)", separate from Search and Training, so a site can welcome search crawlers, refuse training, and decide about agents on its own terms.

Identity solves only half the problem. A signature proves which platform sent a request, not which person the agent works for or what it is allowed to do. That second layer, delegated and scoped agent access, is where identity companies such as Okta, Stytch, and 1Password are now building. Until it matures, the safest pattern for anything behind a login is still an official API or integration.

Publishers have their own side of this. A site can state in robots.txt which crawlers may read it and, with Content Signals, what AI systems may do with what they read. Our guide to AI robots.txt generators covers the publisher view.

How Reliable Are Browser Agents? What the Benchmarks Say

Browser agents are reliable for narrow, repeated tasks and still unreliable for long, open-ended ones. Public benchmarks show strong scores on curated tasks and noticeably lower scores on live websites, and the gap is the lesson: the real web is messier than any test set.

Benchmark What it measures What it showed
OSWorld (2024) Real computer tasks across desktop apps and the web At launch, humans completed 72.36% of tasks and the best model only 12.24%
OSWorld-Verified (2026) The maintained version of the same benchmark By August 2026 the best general model on the official results scored about 86%, and an agent framework about 90%
WebVoyager Tasks on live websites Browser Use reported 89.1% in December 2024, on 586 tasks after removing 55
Online-Mind2Web (2025) Live web tasks graded by humans Frontier agents showed "a drastic drop in success rate" compared with WebVoyager. Under human grading, Browser Use scored 30.0%, Claude computer use 56.3%, and OpenAI's Operator 61.3%.

The gap between WebVoyager and Online-Mind2Web is the useful number. Scores collected on a curated set, or graded by another model, run far higher than scores on fresh live tasks graded by people. When a vendor quotes a success rate, ask which benchmark, which grader, and how many tasks were removed.

Benchmarks also hide the failure that matters most in production: the confident wrong turn. An agent that clicks the wrong "Submit" and reports success is worse than one that stops and asks. That is why builders who run agents at scale converge on the same habits.

  1. Keep steps atomic. "Click the checkout button" is easier to get right, and to check, than "complete the purchase."
  2. Pick the model per task. Browserbase publishes browser-task evals for exactly this reason: a model that tops a coding benchmark can still fail at a date picker, which Klein calls "the bane of the existence of LLMs."
  3. Cache what worked. Replaying a successful action skips the model and removes a chance to go wrong.
  4. Record everything. A session replay turns a mystery failure into a five-minute fix.
  5. Put a person on the first runs. Promote a workflow to a schedule only after it passes repeatedly.

How to Build or Choose a Browser Agent

Choose by who you are. People who want tasks done use an in-browser assistant or a hosted cloud agent. Developers building agents into a product use a framework such as Stagehand or Browser Use on cloud browsers. Teams that mainly need web data should start with search and extraction, not a browser.

If you are... Start with Why
An individual with web chores Claude in Chrome, Perplexity Comet, or the cloud browser in ChatGPT Work No setup, works with your logins
A developer adding web actions to a product Stagehand or Browser Use on a cloud browser Code for the known steps, AI for the flexible ones
A developer whose AI tool needs a browser Playwright MCP locally, or a hosted browser MCP server One config line gives an assistant a browser
A team that needs web data on a schedule Search and page extraction, then an automation Cheaper and more reliable than clicking
A team whose workflow spans many tools APIs and integrations first, a browser step only where none exists Fewer moving parts, fewer failures

What building with Stagehand looks like

Stagehand, the open-source framework from Browserbase, is a good example of the hybrid style most teams settle on. Developers write known steps as code and use three natural-language methods where pages vary: act takes one action, extract pulls data into a schema, and observe lists what can be done on a page. Version 3 dropped the Playwright dependency in October 2025, and version 4 moved the core logic into a browser extension in August 2026.

  Hybrid browser automation, the Stagehand way

known steps → code open the page, set the region, wait for load
flexible steps → natural language
act("click the comments link for the top story")
extract("title, points, and author of the top story", schema)
observe("actions related to checkout")
output → typed data { title, points, author }

The design principle is worth borrowing even if you never use the library. Teams told Browserbase they did not trust AI to choose their steps, but they did trust it to repair a step that broke. So the steps stay in code, and AI handles the parts that change.

Where Taskade Fits: The Workspace Around the Agent

Taskade is not a browser agent. Taskade AI agents search the web and read pages, and Taskade automations add Search Web, Scrape Webpage, and HTTP Request actions plus 100+ bidirectional integrations. That covers the first three rungs of the ladder. For the fourth, teams call a browser-agent service, and Taskade is where the results land.

That split follows the ladder. Most of what teams ask an agent to do on the web is research, monitoring, and moving data between tools, and none of that needs a browser to click anything.

Rung What Taskade does
1. Search AI agents search the web as a built-in tool, and automations use the Search Web action
2. Fetch and extract Agents read pages, and automations use Scrape Webpage and Summarize Website
3. APIs and integrations 100+ bidirectional integrations, plus the HTTP Request action for any API
4. A real browser Not built in. An automation can call a browser-agent service through HTTP Request or the MCP Client action.
search, read,call APIs result Slack, CRM, email,100+ integrations writes back The web Taskade projectMemory Browser-agent servicefor clicks and logins Taskade AI agentIntelligence Taskade automationExecution Your team and customers
search, read,call APIs result Slack, CRM, email,100+ integrations writes back The web Taskade projectMemory Browser-agent servicefor clicks and logins Taskade AI agentIntelligence Taskade automationExecution Your team and customers

What Taskade adds is the part every browser-agent demo skips: a place for the work to live. Results land in a project your team can see in multiple project views. An agent with persistent memory reads them, compares them with last week, and flags what changed. An automation routes them to Slack, a CRM, or an inbox. That loop is Workspace DNA: Memory feeds Intelligence, Intelligence triggers Execution, and Execution creates Memory.

An autonomous Taskade agent loop working through a task

A Taskade agent works through a multi-step task inside the workspace, where every result stays visible to the team.

Scheduling a recurring automation flow in Taskade

Schedule a flow that searches, reads, and files what it finds, no browser required.

For a hands-on version, our guide to AI web scraping without code builds a scheduled agent that reads public pages into a table. Taskade runs on 15+ frontier models from OpenAI, Anthropic, and open-weight providers, with role-based access from Owner to Viewer, and paid plans start at $10 a month billed annually. Start with a free AI agent →

Key Takeaways

  • A browser agent is a model driving a real browser: observe, decide, act, check, repeat.
  • The first consumer wave, Operator, ChatGPT agent, Atlas, and Project Mariner, was folded into bigger products or shut down by August 2026. The capability moved into browsers people already use and into developer infrastructure.
  • Follow API-first, browser-last: search, then fetch, then an API, and only then a browser.
  • Cost scales with steps, not just model price. Keep steps small and cache what works.
  • The big risk is prompt injection. Use scoped accounts, allow-lists, confirmation gates, and credential vaults.
  • The web is moving from stealth to identity: signed agents, Web Bot Auth, and publisher policies.
  • Results need a system of record. The click is the easy part. The workspace around it is what makes it useful.

🔗 Related Reading

  • History of Browserbase: how the leading cloud-browser company for agents was built
  • What Are AI Agents?: the perceive, reason, act, learn loop behind every agent
  • AI Web Scraping Without Code: rungs 1 and 2 of the ladder, on a schedule
  • LLM vs RAG vs AI Agent vs Agentic AI: where browser agents sit on the capability ladder
  • How LLMs Got Hands: The History of Tool Use: from function calling to computer use
  • History of AI Agents: from SHRDLU to the agent loop
  • Best AI Agents in 2026: the agents people actually use
  • Manus AI Review: a general agent built on a virtual computer
  • Best MCP Servers: including browser servers for AI assistants
  • AI Robots.txt Generators: the publisher side of AI crawlers and agents
  • Wiki: Browser agents, Computer-use agents, Agent sandbox, Agent permissions, MCP client

📚 Sources

  • OpenAI Help Center: ChatGPT agent and Evolving Atlas into ChatGPT
  • Google: Project Mariner and the Gemini 2.5 Computer Use announcement
  • Anthropic: Developing computer use
  • Cloudflare: Web Bot Auth, signed agents, pay per crawl, and Browser Run
  • IETF: Web Bot Auth working group
  • Browserbase: Stagehand v3, Stagehand v4, identity, and pricing
  • Latent Space: Open Operator, Serverless Browsers and the Future of Computer-Using Agents (February 2025)
  • 1Password and Browserbase: Secure Agentic Autofill (October 2025)
  • Cover image: Taskade illustration

💬 Frequently Asked Questions About Browser Agents

What is a browser agent?

A browser agent is an AI system that operates a real web browser to finish a task. It reads the page, decides the next step with a language model, and clicks, types, scrolls, or navigates until the goal is met. It reaches websites that have no API, such as portals, forms, and legacy dashboards.

How is a browser agent different from a headless browser?

A headless browser is the engine, a real browser running without a window that follows exact instructions from code. A browser agent is the driver, a language model that decides what to do on each page. Most browser agents run on top of a headless browser, locally or in the cloud.

How is a browser agent different from web scraping?

A scraper reads pages and extracts data, usually without clicking anything. A browser agent acts: it logs in, fills forms, and moves through multi-step flows. Scraping is faster and cheaper for public pages, and a browser agent is for tasks that need interaction.

What is the difference between a browser agent and a computer-use agent?

A browser agent works only inside a web browser. A computer-use agent controls a whole desktop or phone, including native apps and system dialogs, usually from screenshots. Browser agents are cheaper and more reliable for web tasks, and computer-use agents reach software with no web interface.

What happened to OpenAI Operator and ChatGPT agent?

OpenAI folded Operator into ChatGPT agent in July 2025 and closed the standalone Operator site on August 31, 2025. ChatGPT agent was no longer available by July 2026, when OpenAI pointed users to ChatGPT Work and its cloud browser. The Atlas browser stopped working on August 9, 2026.

Which browser agents are still available in 2026?

Claude in Chrome, Perplexity Comet, Chrome's auto browse, and the cloud browser in ChatGPT Work cover most consumer use. Developers build with Stagehand or Browser Use on cloud browsers from Browserbase, Kernel, Anchor Browser, Steel, Hyperbrowser, or Cloudflare Browser Run.

When should I use an API instead of a browser agent?

Whenever one exists. An API call is faster, cheaper, and more reliable, and it does not break when a page changes. Use a browser agent only for steps with no API, integration, or export.

Are browser agents safe?

With guardrails, for many tasks. The main risk is prompt injection, where text on a page tricks the agent. Use scoped accounts, domain allow-lists, confirmation before payments or deletions, credential vaults, and a person watching the first runs.

How much does a browser agent cost to run?

Browser time plus a model call per step, plus retries. A 20-step task means about 20 model calls. Cloud browsers are billed by the hour, for example 100 browser hours in Browserbase's $20 Developer plan, and reading a page is far cheaper than driving one.

How do websites know a visitor is an AI agent?

Agents increasingly sign their requests. Web Bot Auth lets a platform sign each request so a site or CDN can verify it, and Cloudflare launched signed agents on this basis in August 2025. Unsigned automated traffic is more likely to be blocked.

Are browser agents reliable enough for production?

For narrow, repeatable tasks with guardrails, yes. Keep steps small, cache actions that worked, retry failed steps, record sessions, and route failures to a person.

Can Taskade agents browse the web?

Taskade agents search the web and read pages, and automations add Search Web, Scrape Webpage, and HTTP Request actions plus 100+ bidirectional integrations. Taskade does not click or log in on other websites. Teams call a browser-agent service for those steps, and Taskade holds the results.

What is Stagehand?

Stagehand is Browserbase's open-source framework for controlling a browser with code plus natural language, through the methods act, extract, and observe. Version 3 dropped its Playwright dependency, and version 4 moved into a browser extension.

What is the API-first, browser-last rule?

Reach for the web in order: search, then fetch and extract, then an API or integration, and only then a real browser. Each rung costs more and fails more often than the one before it.


The first web browser was built so a person could read and write the web. The new ones are built so software can. The winning pattern in 2026 is not the agent that clicks the most. It is the one that clicks only when it must, and hands everything else to a workspace that remembers. ▲ ■ ●

Build your first AI agent in Taskade →

0%

On this page

What Is a Browser Agent?How Does a Browser Agent Work?How an agent "sees" a pageA browser agent run, step by stepThe Four Kinds of Browser Agents in 2026What Happened to Browser Agents in 2025 and 2026?Why the first wave foldedWhen Should You Use a Browser Agent? The API-First, Browser-Last LadderHow Much Does a Browser Agent Cost?Are Browser Agents Safe? Prompt Injection and PermissionsHow Do Websites Know an Agent Is an Agent?How Reliable Are Browser Agents? What the Benchmarks SayHow to Build or Choose a Browser AgentWhat building with Stagehand looks likeWhere Taskade Fits: The Workspace Around the AgentKey Takeaways🔗 Related Reading📚 Sources💬 Frequently Asked Questions About Browser Agents

Related Articles

Automation steps executing in sequence inside a Taskade workspace, illustrating how work passes from one step to the next
September 23, 2026AI

Agent Handoff Explained: The Narrow Channel Between AI Agents (2026)

Multi-agent systems fail at the seams, not inside the agents. What a handoff carries, the five ways it breaks, and how t...

Connectome extraction procedure diagram showing how brain imaging becomes a network graph. Image: Wikimedia Commons / Hagmann P, Cammoun L, Gigandet X, Meuli R, Honey CJ, et al / CC BY 3.0
September 22, 2026AI

The Complete Connectome Explained: What a Finished Brain Map Tells Us About AI Agents (2026)

Scientists have now mapped an entire animal nervous system, neuron by neuron. What the map shows, what it cannot show, a...

Previewing and customizing a branded AI agent in Taskade before publishing it as a public page, a custom domain, or a website widget
August 31, 2026AI

Generate the Art. Preview the Agent. Put It on Your Domain (2026)

Ship an AI agent that looks like your company: generate its art in the workspace, preview it the way a visitor sees it, ...

An AI agent with its own tools running inside a Taskade automation
August 30, 2026AI

Agentic Automation Explained: Agent vs AI Step (2026)

Agentic automation runs a named agent with memory and tools inside a workflow. Here is the five-point test that separate...

Construction scaffolding around a building site, the visual metaphor for the scaffolding tax in AI prompting: structure you build to support the work, which eventually has to come down
August 30, 2026AI

The Scaffolding Tax: Why Less Prompt Beats More (2026)

Two frontier labs and two research papers reached the same conclusion by 2026: the AI instructions you wrote to improve ...

Richard S. Sutton, author of The Bitter Lesson and 2024 Turing Award winner, photographed in December 2025. Photo: Wikimedia Commons / Xuthoria / CC BY-SA 4.0
August 29, 2026AI

The Bitter Lesson Explained: Richard Sutton's 26 Words (2026)

Richard Sutton compressed The Bitter Lesson into 26 words. What it means, why he now says LLMs break it, and what it cha...

View All Articles
Browser Agents Explained: How AI Agents Use the Web (2026)